Jul 2026· Journal of Analytical Science and Technology· Vol 17· 0 citations· 46 references
TL;DR
MetaRT, a stacked-ensemble machine learning framework designed to predict the RTs from small dataset of hydrophobic peptides, outperformed both base learners and the ensemble models, achieving a lower root mean square error (RMSE) and a maximum RT deviation.
Abstract
Predicting peptide retention time (RT) remains a significant challenge, particularly when training data is limited. In this study, we present MetaRT, a stacked-ensemble machine learning framework designed to predict the RTs from small dataset of hydrophobic peptides. Peptides composed of hydrophobic amino acids—phenylalanine (F), isoleucine (I), methionine (M), and tryptophan (W) were synthesized, and their experimental RTs were measured from the mixture entities. MetaRT utilizes a graph convolutional network (GCN) to extract structural features from the peptide sequences. The MetaRT model architecture employed multiple base learners, integrating the outputs through a meta-learner optimized via hyperparameter tuning and 3-fold cross-validation. Besides, the performance of MetaRT was compared to three ensemble methods - weight averaging, bagging, and boosting. The results demonstrated that structure-based MetaRT outperformed both base learners and the ensemble models, achieving a lower root mean square error (RMSE) of 0.08 and a maximum RT deviation of approximately 1.4 min. Compared to the prediction performance on molecular descriptors inclusion, the structure-guided model consistently performed well in terms of RMSE. Notably, MetaRT accurately predicted the RTs of sequence isomers by leveraging the structural features, with deviations ranging from 0.2 to 1.3 min. In contrast, descriptor-based model showed increased prediction error for the isomeric sequences. For peptides with lower hydrophobicity that were not included in the training data, the structure-based predictions led to the maximum deviation of 4.9 min from the experimental RTs. The entire predicted RTs were subsequently validated by linear regression analyses with the corresponding experimental values. These findings highlight the potential of MetaRT as a structure-based predictive tool for improving RT prediction accuracy, especially in data-limited scenarios. Future work will focus on enhancing the robustness of MetaRT by incorporating a wider variety of peptide classes to further refine its predictive capabilities.
Peptide hormones are important signaling molecules that regulate diverse physiological processes and have substantial therapeutic relevance. However, their experimental identification can be challenging because of low abundance, limited stability, and complex post-translational processing, highlighting the need for rel...
Recently, a new category of machine learning approaches for tabular data has emerged: tabular foundation models (TFM), based on in-context learning. A TFM is a neural network (usually a transformer) pretrained primarily on synthetic data. Its input is an entire data set: features and labels for training records, along...
D. Matyushin, A. Sholokhova· Journal of Chemical Informat...· 0 citations
INTRODUCTION
Experimental identification of anticancer peptides (ACPs) is timeconsuming and costly, which limits large-scale ACP discovery and screening. To address this challenge, we developed MDFA-MLP, a novel computational framework for ACP prediction that integrates multi-scale feature learning and ensemble classif...
Results indicate that combining protein language models with sequence modeling and class-imbalance learning strategies is an effective way to improve peptide hormone prediction and to improve the recognition of hormone peptides as the minority class.
Chun-Yan Ao, Shihu Jiao, Xi Su et al.· Journal of Advanced Research· 0 citations
In recent years, deep-learning has revolutionized protein structure prediction, achieving remarkable speed and accuracy. RNA structure prediction, however, has lagged behind. Although several methods have shown moderate success in predicting RNA secondary and tertiary structures, none have reached the accuracy observed...
Conner J. Langeberg, Taehan Kim, Roma Nagle et al.· RNA: A publication of the RN...· 0 citations
TPpred-PepPA is developed, a two-stage hierarchical deep learning framework based on the ProtT5 pre-trained large language model that achieves state-of-the-art predictive performance and provides valuable interpretability for the discovery of multi-functional therapeutic peptides.
Ke Yan, Si-Yang Lu, Shutao Chen et al.· BMC Biology· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.