Jul 2026· Journal of Chemical Information and Modeling· 1 citation· 33 references
Medicine
TL;DR
This work introduces a method to smoothly transition from physics-based to knowledge-based predictions based on the uncertainty of each model and shows that combining structure-based and ML models significantly improves the prediction accuracy if training data is limited, whereas the weighting smoothly shifts from docking to ML as more data is acquired.
Abstract
Protein-ligand binding affinity prediction is central in computer-aided drug design, and a wide range of physics-based, empirical, and machine learning (ML) tools have been developed for this purpose. Scientists are often tasked with choosing an optimal tool for a given discovery problem, which can be difficult. Physical models typically perform best in the absence of target-specific data but are often outperformed by ML models as the amount of data on a given target grows. Here, we introduce a method to smoothly transition from physics-based to knowledge-based predictions based on the uncertainty of each model. We apply this framework to combine docking scores with predictions from a Gaussian Process model trained on binding affinities from two industrial data sets. We show that combining structure-based and ML models significantly improves the prediction accuracy if training data is limited, whereas the weighting smoothly shifts from docking to ML as more data is acquired. We also show that structure-based methods, being insensitive to the distribution of binders across chemical space, can improve generalizability to new chemistries and increase the hit rate in active learning screens where the binder density varies among scaffolds in a virtual library. Thus, integrating predictions from multiple tools not only optimizes the use of limited experimental data but also ensures more robust performance compared to reliance on a single model.
This study shows that incorporating synthetic molecular dynamics data improves deep learning models for protein–ligand binding affinity prediction beyond static experimental structures, and highlights that dynamic synthetic datasets can enable deep learning models to outperform conventional methods such as MM-PBSA while remaining computationally efficient.
P. Agrawal, Prathit Chatterjee, U. Priyakumar· Journal of Cheminformatics· 0 citations
The resulting model, HydrAffinity, is an interaction-free, dynamic sparse model that uses pre-trained encoders and MoE for parameter-efficient learning and outperforms all interaction-free methods and matches state-of-the-art interaction-based methods on CASF-2016.
An extensive evaluation of Boltz-2 using two large-scale data sets shows that Boltz-2 lacks the energetic resolution required for lead identification, highlighting the necessity of employing physics-based methods for the reliability and refinement of AI-derived models.
S. Wan, Xibei Zhang, Xiao Xue et al.· Journal of Chemical Theory a...· 1 citation
Boltz is benchmarked using a curated set of ligand-bound human G protein-coupled receptors from families unseen during training, showing that while Boltz generally predicts receptor backbones accurately, ligand poses can contain significant errors that lead to a limited ability to reproduce experimental affinity data when tested with FEP+.
Lichirui Zhang, R. Friesner, Edward B. Miller et al.· npj Drug Discovery· 0 citations
In the early stages of drug discovery, predicting drug-target affinity is a crucial task. Due to the vast scale of genomic and chemical spaces, traditional biological methods are time-consuming, labor-intensive, and resource-demanding. As a result, machine learning-based computational methods have emerged to narrow down the pool of drug candidates. However, machine learning approaches still face several challenges in practical applications, particularly the scarcity of labeled samples and poor model generalization capability. To address these issues, this paper proposes a novel drug-target affinity prediction model, termed MetaBayes-DTA, based on an uncertainty-aware meta-learning framework. The model integrates the few-shot rapid adaptation capability of meta-learning with an uncertainty quantification mechanism to enhance prediction accuracy and reliability. MetaBayes-DTA is evaluated on two benchmark datasets, DAVIS and KIBA. Experimental results demonstrate that the proposed model outperforms existing methods.
Naihan Shi, Yanpeng Zhao, Wanying Li et al.· 2026 IEEE 27th China Confere...· 0 citations
CMD-PLA is proposed as a dynamics-aware framework for protein-ligand affinity prediction, and the conformational evolution of the ligand is explicitly modeled as a pocket-dependent dynamical process rather than an isolated static update.
Hao Li, Dongjiang Niu, Xiaofeng Wang et al.· Computational biology and ch...· 0 citations