It is demonstrated that CV R² computed from as few as 50 labeled peptides can be sufficient to estimate final active learning end-point performance, providing a practical, data-efficient criterion for deciding whether a given dataset combined with SCARSE is suitable for iterative peptide discovery.
Abstract
Reliable estimation of downstream performance in low-data peptide machine learning is critical for guiding early-stage AI-driven peptide engineering. Yet, it is often unclear how to assess whether a model will be effective in iterative discovery settings. Here, we show that the cross validation R² score can serve as a simple and robust proxy for predicting active learning workflow performance, enabling early-stage evaluation of model suitability for sequential peptide optimization. To support this, we introduce SCARSE, a machine learning framework combining ESM-2 protein language model embeddings with Gaussian process regression and extremely randomized trees classification, designed for low-resource peptide property prediction (20–500 training samples). We benchmark SCARSE across 23 peptide and small-protein datasets covering substitution and indel variants, antimicrobial peptides, cell-penetrating peptides, and toxic/non-toxic peptides. SCARSE significantly outperforms a hand-engineered descriptor baseline on substitution and indel tasks, while comparable performance was achieved on shorter peptide non-mutant datasets where simpler descriptors capture enough of the signal. In simulated active learning workflows, SCARSE consistently outperforms baseline and random sampling strategies. Notably, we demonstrate that CV R² computed from as few as 50 labeled peptides can be sufficient to estimate final active learning end-point performance, providing a practical, data-efficient criterion for deciding whether a given dataset combined with SCARSE is suitable for iterative peptide discovery. SCARSE is released as a pip package and is available via HuggingFace Spaces to facilitate integration into peptide engineering workflows.
Machine learning (ML) has accelerated molecular discovery, yet training models to generalize to out-of-distribution (OOD) chemical spaces remains fundamentally constrained by the high cost of experimental validation. In antibiotic discovery, where whole-cell phenotypic high throughput screening (HTS) is resource-intensive, iterative ML-guided compound selection – or Active Learning (AL) – offers a pathway to efficiently navigate available chemical spaces. However, the algorithmic tradeoffs between prioritizing compound novelty (exploration), predicted bioactivity (exploitation), and their impact on OOD generalizability remain unresolved for noisy, whole-cell biological systems. In this work, we systematically evaluate three AL strategies for whole-cell bacterial bioactivity and benchmark their effects on model accuracy, hit rate, and OOD performance. Using retrospective simulations on Mycobacterium tuberculosis HTS data, we identify an optimal AL strategy that balances predicted hit/non-hit novelty with overall hit rate. We then integrate the strategy in a closed-loop Borrelia burgdorferi antibiotic discovery HTS campaign. The AL-guided approach successfully increased the experimental screening hit rate five-fold (from a 0.2% rate within investigator-selected plates to 1.0%). Further, when the trained model was applied in prospective in silico selection of highly diverse compounds across multiple bacterial species, the AL-trained whole-cell inhibition predictor demonstrates 53-fold enrichment over investigator-directed screening (11.0% experimental validation of predicted hits). Of these, 100% demonstrated the intended narrow spectrum activity for Borrelia burgdorferi. These results demonstrate that calibrated AL strategies can overcome data acquisition bottlenecks and train generalizable property predictors able to extrapolate to OOD molecules.
Lia R. Serrano, Andrew Zhou, Ziming Wei et al.· bioRxiv· 0 citations
The proposed two-phase generative–evolutionary framework provides a generalizable approach for balancing functional optimization and distributional realism and can be applied to peptide discovery and data augmentation in imbalanced biological datasets thereby generating high confidence peptides for wet lab validation.
This review systematically examines the key methodological innovations, including peptide representation learning, multi-modal fusion strategies, multi-label learning paradigms, and emerging predictive frameworks empowered by deep neural architectures and ProtLM-based embeddings, and summarizes the practical applications of these models in peptide database mining, functional mechanism interpretation, and mutation effect prediction.
An unsupervised machine-learning framework that leverages hybrid high-dimensional peptide representations to discover high-performance AFPT families without requiring 3D structures or large labeled data sets is presented and demonstrates how unsupervised hybrid-feature learning can reveal actionable biophysical design rules from sequence data alone.
Nazmul Shuzan, Jialun Wei, Jie Zheng· Journal of Chemical Informat...· 0 citations
A data-driven, multi-objective peptide design framework that inte-grates sequence-to-feature transformations using Fast Fourier Transform - based representations, and metric-learning based optimization strategies, to provide an interpretable and computationally efficient alternative for peptide design under limited-data constraints.
A. Trinh· Proceedings of the 3rd Found...· 0 citations
Cutting-edge bioinformatics research is increasingly intertwined with pre-trained model techniques. However, achieving superior performance of these models in downstream applications typically requires large amounts of accurately labeled experimental data for fine-tuning, which poses substantial practical challenges due to the difficulty in preparing such datasets at scale. To address this limitation, we propose a novel few-shot fine-tuning framework, the Evolution-Aware Adaptation (Evo-AA). It aligns fine-tuning with pre-training objectives while integrating prompt learning and biological coevolutionary insights. Additionally, we introduce reinforced prompting and lambda ranking loss to further improve performance. Extensive experiments demonstrate that Evo-AA with limited training set, enhances the spearman correlation in fine-tuning tasks, while achieving superior precision and recall rates in homology search tasks. Our findings suggest that Evo-AA holds great potential to drive advancements in protein engineering and computational biology.
Yuxuan Wu, Huiqun Yu, Guisheng Fan et al.· Annual International Compute...· 0 citations