This review surveys state-of-the-art methods across drug–target interaction prediction, protein–ligand complex modeling and docking, de novo molecular generation, and biomolecule design, examining the convergence of docking, structure prediction, and molecular generation within co-folding and diffusion-based frameworks.
Abstract
Abstract Recent advances in protein structure determination and prediction, large-scale structural databases, and artificial intelligence have reshaped structure-based drug discovery. Structure-aware artificial intelligence models integrate molecular representation learning with three-dimensional protein information to model interactions, predict complex structures and binding poses, and generate novel molecules. In this review, we follow this paradigm along a continuum from protein–ligand modeling to the de novo design of biomolecular binders. We first outline the molecular and protein representations that render structures computable, together with the growing collection of structural data resources. We then survey state-of-the-art methods across drug–target interaction prediction, protein–ligand complex modeling and docking, de novo molecular generation, and biomolecule design, examining the convergence of docking, structure prediction, and molecular generation within co-folding and diffusion-based frameworks. Despite these advances, prospective experimental validation remains scarce, and persistent limitations such as biased structural coverage, limited and ambiguous negative supervision, fragmented benchmarking, and insufficient mechanistic interpretability continue to constrain real-world utility. Progress in data quality and supervision design, evaluation rigor, and design-relevant interpretability will be essential to translate methodological innovation into practical impact.
Accurately modeling biomolecular interactions is a central bottleneck in biology and therapeutic discovery. Here, we introduce Open Drug Discovery Engine (OpenDDE), an open-source, all-atom biomolecular foundation model that uses co-folding as the entry point to a scalable AI-driven drug discovery engine. Rather than treating structure prediction as an isolated endpoint, OpenDDE is designed as a shared structural reasoning layer for modeling sequence-structure-function relationships across biomolecular complexes, enabling complex structure prediction today while providing a foundation for de novo design, affinity estimation, structure-conditioned optimization, and more. OpenDDE integrates advances in all-atom architecture, atomic latent reasoning, inference optimization, and large-scale data processing to achieve IsoDDE-level co-folding accuracy within a reproducible and openly accessible framework. We also identify two scaling-law directions for co-folding models, revealing practical routes for continued improvement through data, model, inference, and training scaling. By releasing training code, inference pipelines, checkpoints, and benchmarks, OpenDDE aims to democratize access to frontier biomolecular intelligence, accelerate global collaboration, and lay an open foundation for next-generation drug discovery systems that can move from predicting molecular structures toward designing, scoring, and optimizing therapeutic candidates for human health.
Vilya-2 is the structure-prediction oracle that de novo peptide design pipelines require--establishing the all-atom approach as a general foundation for the design and evaluation of de novo peptide therapeutics.
Vilya Research Pascal Sturmfels, Naozumi Hiranuma, M. Salem et al.· 0 citations
Natural products (NPs) have historically yielded numerous therapeutic agents, yet their integration into modern drug discovery has been constrained by chemical complexity, low abundance, laborious dereplication, and limited target annotation. Convergence of multi-omics technologies with high-resolution structural and biological data has created unprecedented opportunities for artificial intelligence (AI) to accelerate NP-based therapeutics development. This review provides an operational, end-to-end workflow that explicitly connects computational predictions to medicinal chemistry decision points, addressing a critical gap between computational prediction and clinical translation. We trace the complete discovery pipeline: computational mining of biosynthetic gene clusters (BGCs) and metabolomes, deep learning (DL)-assisted structural elucidation and dereplication, network-based target identification using protein-ligand prediction, and generative molecular design inspired by NP scaffolds (including large language models, diffusion models, and genetic algorithms). Critical evaluation of current limitations (data scarcity, lack of standardized ontologies, model interpretability) is complemented by discussion of emergent strategies (foundation models trained on multi-modal data, graph neural networks, autonomous closed-loop laboratories). Representative case studies, including the synthetic AI-designed clinical benchmark rentosertib, illustrate the current evidence spectrum from discovery-level validation to early clinical benchmarking, while also highlighting that most AI-enabled NP discovery workflows remain at the preclinical or proof-of-concept stage, with limited quantitative evidence of improved clinical productivity. We conclude with an Outlook proposing feasible developments for 2025-2030: self-driving laboratories with reported acceleration in specific experimental contexts, foundation models enabling hypothesis-free chemical space exploration, and sustainability-aware AI frameworks embedding biodiversity impact assessments. This operational focus fills a critical gap between algorithmic capability and clinically actionable NP-derived leads. Importantly, while AI has demonstrably accelerated several early discovery steps, quantitative comparisons with classical NP workflows remain limited, and most reported advances are supported by preclinical or proof-of-concept studies rather than systematic evidence of improved time-to-lead, cost reduction, or clinical success rates.
Antonio Lavecchia· Medicinal research reviews (...· 0 citations
Molecular docking and molecular dynamics (MD) simulations have become indispensable tools in modern drug discovery, enabling researchers to accelerate the identification and optimisation of therapeutic compounds. This comprehensive review examines the fundamental principles underlying these computational approaches, their diverse applications in pharmaceutical development, and the significant limitations that currently constrain their predictive accuracy and applicability. We discuss structure-based drug design methodologies, scoring functions, binding-affinity prediction, conformational sampling strategies, and the integration of artificial intelligence into computational drug discovery. Furthermore, we address critical challenges, including protein flexibility representation, ADMET (Absorption, Distribution, Metabolism, Excretion, and Toxicity) prediction accuracy, and the persistent discrepancy between in silico predictions and experimental validation. Recent advances in hardware acceleration, force-field development, and machine learning are reshaping the landscape of computational drug discovery. This review synthesises current knowledge and highlights future opportunities for enhancing the reliability and efficiency of molecular docking and simulation studies in pharmaceutical research.
Perli.Kranti Kumar, S. Nilewar· Indian Journal of Pharmaceut...· 0 citations