The proposed two-phase generative–evolutionary framework provides a generalizable approach for balancing functional optimization and distributional realism and can be applied to peptide discovery and data augmentation in imbalanced biological datasets thereby generating high confidence peptides for wet lab validation.
Abstract
Recent advances in artificial intelligence have accelerated the discovery of bioactive peptides by enabling computational exploration of the vast peptide sequence space. However, existing peptide generation approaches generally rely on either distribution-learning models, which generate biologically realistic sequences but do not consistently optimize functional activity, or optimization-based methods, which maximize prediction confidence while often deviating from the underlying distribution of experimentally validated peptides. To address this limitation, a two-phase generative–evolutionary framework is proposed that integrates distribution learning with evolutionary optimization. In the first phase, Variational Autoencoders (VAE), Autoregressive Transformers (ART), and Token Diffusion Transformers (TDT) are used to generate biologically plausible seed peptides. In the second phase, these peptides were used as initial seed for Hill Climbing optimization procedure that iteratively improves fitness function score. The proposed two-phase framework was evaluated using a dataset of experimentally validated IL-2-inducing peptides. Evaluation using independent IL-2 prediction models showed that Autoregressive Transformer combined with Hill Climbing achieved the best overall performance, achieving the mean IL-2 induction confidence score of 0.96 while reducing KL divergence from 2.26 for standalone Hill Climbing to 0.75. A case study on an independent IL-13 inducing peptide dataset showed similar trends, with ART initialized Hill Climbing achieving the mean IL-13 induction score of 0.99 while reducing KL divergence from 1.76 to 0.59. Overall, the framework provides a generalizable approach for balancing functional optimization and distributional realism and can be applied to peptide discovery and data augmentation in imbalanced biological datasets thereby generating high confidence peptides for wet lab validation. Highlights Proposed a two-phase framework for bioactive peptide generation with potential to address class imbalance in peptide classification tasks. Performed a systematic comparison of distribution-learning and optimization-based approaches for peptide generation. Combined distribution-learning models for sequence generation with optimization algorithms for improving peptide functional properties. Demonstrated the applicability of the proposed framework across multiple bioactive peptide datasets.
This review systematically examines the key methodological innovations, including peptide representation learning, multi-modal fusion strategies, multi-label learning paradigms, and emerging predictive frameworks empowered by deep neural architectures and ProtLM-based embeddings, and summarizes the practical applications of these models in peptide database mining, functional mechanism interpretation, and mutation effect prediction.
A snapshot of AI-driven technologies for AMP design is provided and two modes of AI-driven technologies for AMP design are surveyed, one concentrated on identifying whether current data possess antimicrobial activity and the other on generating AMP candidates with potential therapeutic properties (generation-oriented).
Yongqiang Liu, Jie Hu, Ning Zhang et al.· Synthetic and Systems Biotec...· 0 citations
Cyclic peptides have emerged as a compelling class of bioactive scaffolds, but de novo design of target-binding cyclic peptides from protein structures remains challenging. Here, we present HighMorph, an interaction-guided framework that combines protein–protein interaction information with artificial intelligence for rational cyclic peptide design. HighMorph integrates Monte Carlo tree search with a Transformer-based policy-value network to efficiently explore cyclic peptide sequence space, while incorporating explicit atomic-level hydrogen bond constraints extracted from reference protein–protein complexes to guide sequence optimization. The framework is systematically validated on two clinically relevant targets, programmed death-ligand 1 (PD-L1) and kallikrein-related peptidase 4 (KLK4). Notably, 33.3% and 40% of the generated candidates are active against PD-L1 and KLK4, respectively, with active cyclic peptides exhibiting micromolar binding affinities (approximately 10–6 M). These results validate our approach for cyclic peptide design. Additionally, interaction analysis provides insights for developing therapeutics targeting challenging protein interfaces.
Minhui Lan, Chengyun Zhang, Wentong Wang et al.· Journal of Medicinal Chemist...· 0 citations
Targeted peptide therapeutics offer a potent solution for undruggable intracellular targets yet their clinical translation remains hampered by poor membrane permeability and metabolic instability. The integration of high-performance computing and artificial intelligence is currently driving a fundamental transition from empirical screening to rational de novo design. This review moves beyond a conventional enumeration of tools to construct a strategic framework that integrates physics-based validation with generative deep learning. We critically analyze the synergistic application of molecular dynamics and docking for thermodynamic verification while simultaneously evaluating how diffusion models and protein language models accelerate the exploration of vast chemical spaces. By delineating a closed-loop workflow that incorporates pharmacokinetic constraints into generative algorithms this review not only synthesizes current advancements but also provides a strategic roadmap for seamlessly integrating generative AI with physics-based validation, thereby accelerating the transition of computationally designed peptides from in silico blueprints to viable clinical candidates.
Wenjing Hu, Yuantao Sun, Ting Li et al.· The Innovation Drug Discover...· 2 citations
ScrambleBench provides a holistic medicinal chemistry-oriented framework that identifies methodological strengths, limitations, and opportunities for future model development and highlights the importance of evaluating chemical diversity explicitly and using the recently proposed metrics such as Hamiltonian Diversity (HamDiv) which assess both quantity and dissimilarity of a molecular set.
Veincent Yap, Pan Xu, Frankie S. Mak et al.· Journal of Cheminformatics· 0 citations