Aug 2026· BMC Biology· Vol 24· 0 citations· 64 references
Medicine
TL;DR
TPpred-PepPA is developed, a two-stage hierarchical deep learning framework based on the ProtT5 pre-trained large language model that achieves state-of-the-art predictive performance and provides valuable interpretability for the discovery of multi-functional therapeutic peptides.
Abstract
Therapeutic peptides exert pivotal effects in diverse biological processes, and have attracted significant interest in the field of biomedicine in recent years. However, most existing methods often fail to adequately capture the intricate interactions among amino acid residues and the contextual dependencies within peptide sequences, which hampers the extraction of deep semantic representations and ultimately restricts predictive performance. Moreover, the task of multi-functional therapeutic peptide prediction is inherently constrained by the challenge of imbalanced multi-label classification resulting from long-tailed distribution patterns. In this study, we propose a two-stage hierarchical deep learning framework, named TPpred-PepPA, for the prediction of multi-functional therapeutic peptides based on pragmatic analysis. Specifically, ProtT5 is employed to extract deep semantic representations that capture residue-level contextual information. In the first stage, a transformer-based network is utilized to perform shared representation learning, wherein the encoder model captures the intricate inter-residue interaction to characterize the contextual semantics of peptide sequences. In the second stage, the framework is fine-tuned by incorporating task-specific classifiers and optimizing the classification decision with Asymmetric Loss. Then the dynamic thresholding strategy is utilized to address the long-tail distribution problem, enabling more accurate prediction performance of multi-functional therapeutic peptide. Moreover, we adopted the SHAP analysis and motif identification to interpret feature contributions and identify key functional peptide fragments, respectively. Our experimental results indicate that TPpred-PepPA significantly outperforms all current baseline methods in identifying multi-functional therapeutic peptides and exhibits robust performance in recognizing rare functional categories. We developed TPpred-PepPA, a two-stage hierarchical deep learning framework based on the ProtT5 pre-trained large language model. Compared with existing methods, TPpred-PepPA achieves state-of-the-art predictive performance and provides valuable interpretability for the discovery of multi-functional therapeutic peptides. Finally, a web server has been established and is accessible at http://bliulab.net/TPpred-PepPA.
It is shown that single-sequence PLMs can perform in-context peptide learning without gradient updates, task-specific retraining, or architectural modification, and MPEP conditioning is established as a lightweight strategy for low-data peptide classification.
Joshua Almonte, M. Vu, Andrew Ahn et al.· bioRxiv· 0 citations
This paper proposes TextDTI, a multimodal framework that simultaneously exploits sequential and structural representations and enhances feature alignment through adversarial learning and contrastive loss, resulting in robust and high-performance DTI prediction.
Jia-Qi Deng, Senyu Tang, Ji-Jun Tang et al.· Journal of Chemical Informat...· 0 citations
PMAVP is proposed, a multi-task learning framework that integrates the ProtT5 pre-trained protein language model with a Mamba-inspired module for AVP identification and functional activity prediction and introduces Focal Loss to mitigate class imbalance and leverage transfer learning to enhance performance on functiona...
Pei-Wei Wei, Wei-Hao Su, Qing-Song Qin et al.· International Journal of Mol...· 0 citations
Pep-Instructions is established as a unified benchmark and the value of peptide-specific instruction tuning for peptide understanding, prediction, and design is demonstrated.
Kai Yang, Tian-Xiang Wu, Wenbo Zhang et al.· Bioinformatics· 0 citations
MetaRT, a stacked-ensemble machine learning framework designed to predict the RTs from small dataset of hydrophobic peptides, outperformed both base learners and the ensemble models, achieving a lower root mean square error (RMSE) and a maximum RT deviation.
R. A. Mahmood, M. Mahin, R. Alam et al.· Journal of Analytical Scienc...· 0 citations
Overall, DrugPLMFormer provides a reproducible, leakage-aware framework for retrospective sequence-based druggability screening and target prioritization, while prospective validation and experimental confirmation remain necessary before operational deployment.
Z. Kafi, Khosro Rezaee, Hossein Eslami· Journal of King Saud Univers...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.