Aug 2026· Journal of Chemical Information and Modeling· 0 citations· 69 references
TL;DR
An unsupervised machine-learning framework that leverages hybrid high-dimensional peptide representations to discover high-performance AFPT families without requiring 3D structures or large labeled data sets is presented and demonstrates how unsupervised hybrid-feature learning can reveal actionable biophysical design rules from sequence data alone.
Abstract
Antifreeze peptides (AFPTs) offer a potentially nontoxic, sequence-programmable alternative to conventional cryoprotectants for preserving biological materials, yet poorly defined sequence–activity relationships continue to limit rational design. Natural AFPTs are often weak, scarce, or context-dependent, and existing design strategies rely on incremental motif tuning with low hit rates and limited interpretability. Here, we present an unsupervised machine-learning framework that leverages hybrid high-dimensional peptide representations to discover high-performance AFPT families without requiring 3D structures or large labeled data sets. We curated the largest annotated AFPT benchmark to date (n = 719) and embedded sequences in a feature space combining physicochemical descriptors with protein language model (PLM) embeddings. Unsupervised clustering resolved distinct active families within the sequence landscape, quantitatively validated by a subset of 107 peptides with measured single-crystal ice-growth rates. Mechanistic interrogation uncovered a dual-signal architecture not previously codified at the peptide level: (i) a regularly spaced polar ice-binding face encoded by primary-sequence motifs, coupled with (ii) a rigid, glycine-depleted scaffold captured only by latent PLM features. A logistic regression classifier trained on the minimal 10-feature set achieved near-perfect separability of active versus inactive families (AUC = 0.98), confirming the generality of the dual-signal rule. Guided by this interpretable blueprint, we designed 14 de novo peptides─10 dual-signal positive designs and 4 negative controls─and validated them alongside 3 literature-reported benchmarks through multiple orthogonal assays. As predicted, negative controls showed minimal activity across all metrics, whereas dual-signal designs exhibited strong ice recrystallization inhibition (IRI activity up to ∼55%), substantial temperature (down to −3.45 °C), and high red-blood-cell post-thaw recovery (85–92%). This work establishes a generalizable, presynthesis prioritization framework for antifreeze peptide engineering and demonstrates how unsupervised hybrid-feature learning can reveal actionable biophysical design rules from sequence data alone.
Deep learning-based protein design methods are used to redesign the globular fish antifreeze protein AFPIII, keeping the previously reported ice-binding residues fixed, highlighting the value of deep learning-based protein design methods both for generating AFP variants with desirable properties and for uncovering gaps in existing knowledge of well-characterized AFPs.
Cianna N. Calia, Arthur J. Altunc, Rosemary J. Eufemio et al.· bioRxiv· 0 citations
This review systematically examines the key methodological innovations, including peptide representation learning, multi-modal fusion strategies, multi-label learning paradigms, and emerging predictive frameworks empowered by deep neural architectures and ProtLM-based embeddings, and summarizes the practical applications of these models in peptide database mining, functional mechanism interpretation, and mutation effect prediction.
How advances in artificial intelligence and computational modeling may reshape the rational design of next-generation peptide therapeutics is explored and an integrated experimental–computational framework is proposed to facilitate the development of clinically actionable candidates is proposed.
Ha Thi Ngoc Nguyen, B. Le, Nhung Thi Hong Van et al.· Pharmaceuticals· 0 citations
It is demonstrated that CV R² computed from as few as 50 labeled peptides can be sufficient to estimate final active learning end-point performance, providing a practical, data-efficient criterion for deciding whether a given dataset combined with SCARSE is suitable for iterative peptide discovery.
Leo Andrekson, Robin Rydbergh, Rocío Mercado et al.· bioRxiv· 0 citations
DeepAden achieves competitive performance compared with state-of-the-art tools on a benchmark dataset, and enabled the identification of two Streptomyces NRPS gene clusters through accurate A-domain substrates specificity predictions.
A data-driven, multi-objective peptide design framework that inte-grates sequence-to-feature transformations using Fast Fourier Transform - based representations, and metric-learning based optimization strategies, to provide an interpretable and computationally efficient alternative for peptide design under limited-data constraints.
A. Trinh· Proceedings of the 3rd Found...· 0 citations