Skip to content

PepOSX-AI: CPP - an interpretable transformer-based deep learning model for prediction of cell-penetrating peptides.

Jul 2026 · Analytica Chimica Acta · Vol 1418, pp. 345991 · 0 citations · 43 references
Medicine

TL;DR

PepOSX-AI: CPP provides a novel computational tool and theoretical basis for the screening of CPPs and effectively captured long-range dependencies in amino acid sequences through a self-attention mechanism, enabling the interpretability of attention patterns.

Abstract

Background

Cell-penetrating peptides (CPPs) are short-chain molecules capable of enhancing the transmembrane delivery of bioactive substances, displaying extensive application potential in the delivery of functional components and the improvement of their bioavailability. Traditional CPP discovery methods, however, rely on a tedious, step-by-step screening process involving cell and animal experiments, which is highly inefficient.

Methods

The deep-learning model, PepOSX-AI: CPP, was developed based on the Transformer architecture and integrative features of six physicochemical descriptors. This involved a series of explorations of model parameters and automated hyperparameter optimization. The model achieved a high area under the curve (AUC) of 0.914 on the dataset. Furthermore, the model effectively captured long-range dependencies in amino acid sequences through a self-attention mechanism, enabling the interpretability of attention patterns. This capability facilitates the identification of key sequence features influencing peptide penetration ability.

Significance

AND NOVELTY In comparison to some existing CPP prediction models, PepOSX-AI: CPP demonstrated an accuracy of 91.00% on the application test set, highlighting its acceptable predictive performance. It provides a novel computational tool and theoretical basis for the screening of CPPs. Based on these findings, an online platform was also developed to facilitate user application.

View source

Similar papers

Aug 2026

MDFA-MLP: Multi-scale Dilated Fusion Attention and Ensemble Framework for Anticancer Peptide Prediction.

INTRODUCTION Experimental identification of anticancer peptides (ACPs) is timeconsuming and costly, which limits large-scale ACP discovery and screening. To address this challenge, we developed MDFA-MLP, a novel computational framework for ACP prediction that integrates multi-scale feature learning and ensemble classification strategies. METHODS The proposed framework combines physicochemical descriptors, including amino acid composition (AAC), dipeptide composition (DPC), composition-transition-distribution (CTD), and pseudo-amino acid composition (PseAAC), with ProtBERT-derived embeddings. A Multiscale Dilated Fusion Attention (MDFA) module was designed to capture sequence patterns at different scales and enhance feature fusion. An ensemble classifier consisting of a multilayer perceptron (MLP), support vector machine (SVM), and histogram-based gradient boosting (HGB) was employed to improve prediction robustness and stability. RESULTS The proposed model was evaluated on the AntiCP 2.0 dataset under the different negative-sample settings. On Dataset A, MDFA-MLP achieved an accuracy of 93.9%, sensitivity of 92.2%, specificity of 95.8%, and an MCC of 0.89. On the more challenging Dataset B, the model achieved an accuracy of 78.2%, sensitivity of 76.5%, specificity of 82.6%, and an MCC of 0.65. Comparative experiments demonstrated that MDFA-MLP achieved competitive and balanced performance across multiple evaluation metrics. Although the improvement over existing methods was moderate in some cases, the model maintained stable predictive performance under different negative-sample settings, indicating good robustness and generalization ability. DISCUSSION The results indicate that traditional sequence descriptors and deep protein language model embeddings provide complementary biological information. The MDFA module effectively enhances feature representation by integrating multi-scale sequence characteristics, while the ensemble strategy improves model robustness and generalization under varying data distributions. CONCLUSION MDFA-MLP provides an effective and reliable framework for ACP prediction. By integrating handcrafted descriptors, protein language model representations, and ensemble learning, the proposed method can facilitate large-scale computational screening of candidate ACPs prior to experimental validation.

Binyu Li, Zhihua Huang, Yijie Ding et al. · 0 citations
Open access Aug 2026

Artificial intelligence-driven prediction and design of cell-penetrating peptides for advanced drug delivery system

Cell-penetrating peptides (CPPs) are promising delivery vectors for transporting therapeutic agents across cellular membranes. However, their rational design remains challenging because the relationship between peptide sequence and translocation efficiency is highly complex and nonlinear. In this study, we developed an artificial intelligence-driven framework that integrates biochemical rules derived from large language models (LLMs) with conventional peptide descriptors for CPP prediction and design. Interpretable rules extracted from GPT-4o and DeepSeek were encoded as binary feature vectors and combined with sequence-based descriptors to construct hybrid machine learning models. Model performance was evaluated on the benchmark CPP924 dataset using repeated stratified cross-validation, and the optimized models were further used for de novo CPP generation. The resulting candidates were subsequently assessed using multiple established computational benchmarks. The top-performing hybrid classifier achieved a cross-validated accuracy of 0.91 ± 0.03 (best single held-out split, 0.94) on the CPP924 dataset. The LLM-derived rules outperformed conventional physicochemical and fingerprint descriptors and matched amino-acid composition; integrating the rule and composition features yielded the best overall classifier. Using the optimized RF-GPT-Fre and RF-DS-Fre models, we generated six de novo CPP candidates that are sequence-novel (≤53% identity to any training peptide) and retain CPP-like composition and structural features. In-silico evaluation across established tools supports computational prioritisation of these candidates for experimental testing. These findings demonstrate that combining LLM-derived biochemical knowledge with machine learning improves interpretable CPP prediction and candidate prioritisation. This study provides a reproducible computational strategy for peptide engineering and establishes a basis for the experimental evaluation of next-generation drug-delivery vehicles.

Bo Yu, Yue Lu, Ze-Ying Kuang et al. · 0 citations
Open access Aug 2026

A distribution-aware and functionally relevant novel framework for generation and discovery of bioactive peptides

The proposed two-phase generative–evolutionary framework provides a generalizable approach for balancing functional optimization and distributional realism and can be applied to peptide discovery and data augmentation in imbalanced biological datasets thereby generating high confidence peptides for wet lab validation.

Rachit Abhigyan, Vikas Sood, Pooja Arora et al. · 0 citations
Aug 2026

MMU-DPI: Enhancing Generalization in Drug–Protein Interaction Prediction through Multimodal Learning and a Label Mix Strategy

Accurate prediction of drug–protein interactions (DPIs) is crucial for accelerating the drug discovery process. However, the scarcity of experimentally validated interactions can limit the learning of transferable interaction patterns, particularly for previously unseen drugs and proteins. To address this fundamental challenge, we propose the MMU-DPI framework. A key component of this framework is a Label Mix strategy tailored to multimodal DPI prediction, which performs interpolation only in the label space while keeping the input modalities unchanged. This strategy provides stochastic soft-target regularization and improves generalization performance under reduced-data and independent external Cold-both evaluation settings. To effectively process and utilize multimodal data, MMU-DPI adopts a multimodal dual-branch architecture. The first branch uses a Message Passing Neural Network (MPNN) to extract structured representations from drug molecular graphs. It also uses a Convolutional Neural Network (CNN) to capture key biological and functional features from amino acid sequences. The second branch constructs a heterogeneous interaction graph and uses a Graph Attention Network (GAT) to learn deep contextual relationships between drugs and proteins. A learnable global fusion weight combines complementary branch logits to generate the final prediction for each drug–protein pair. Experimental results on multiple benchmark data sets demonstrate that MMU-DPI outperforms several state-of-the-art DPI prediction methods. Case studies further support the ability of MMU-DPI to identify potential DPIs. These results indicate that MMU-DPI can serve as a useful computational tool for drug discovery.

Jiahao Wei, Tie Shen · 0 citations
Jul 2026

PeptideSGCL: Structure-Enhanced Graph-Transformer Encoding and Dual-Level Contrastive Learning for Peptide Property Prediction.

Peptides play important roles in biological processes and biomedical applications, and their hemolytic (Hemo) and nonfouling (NF) properties directly affect their safety and translational potential. Therefore, accurate predictive models are essential for the rational design of functional peptides. Although existing multimodal peptide property prediction methods can jointly exploit sequence and structural information, their structural encoders still rely primarily on local graph convolution and their contrastive objectives are largely focused on cross-modal alignment. Consequently, they remain limited in modeling long-range structural dependencies and in enhancing intramodal discriminability. To address these limitations, we propose a multimodal dual-contrastive learning framework for peptide property prediction, which improves both the structural encoder and the contrastive learning strategy to enhance the quality of joint sequence-structure representations. Specifically, ProtBERT is adopted as the sequence encoder, and a hierarchical GNN-Transformer structural encoder is constructed to capture local topological patterns and long-range structural dependencies. In addition, a parallel graph spatial channel attention module is introduced to enhance task-relevant structural features. Within a shared embedding space, we further design an interintra hybrid supervised contrastive learning strategy to jointly optimize sequence-structure alignment and intramodal class discriminability. Experimental results show that the proposed method achieves overall performance superior to baseline models on both hemolysis and NF prediction tasks, providing an effective framework for multimodal representation learning in peptide-property prediction.

Jiajie Cai, Shuwen Xiong, Yuntao Yang et al. · 0 citations