It is indicated that although many methods report strong performance on standard benchmarks, their effectiveness is often influenced by dataset bias and limited evaluation settings, and most methods exhibit reduced performance in cold-start scenarios, highlighting challenges in generalization.
Abstract
Computational approaches to drug discovery involve multiple sub-problems, and among them, drug-target binding affinity prediction plays an important role. Despite recent advances, accurately predicting binding affinity remains an open research area. The major objective of our paper is to perform a comprehensive review and comparative analysis of recent machine learning methods for drug-target binding affinity prediction, with a focus on identifying strengths, limitations, and research gaps. We review representative recent deep learning approaches that use common benchmark datasets and evaluation metrics, covering a range of neural network architectures and representation strategies. In addition, we analyze seven widely used benchmark datasets and commonly adopted evaluation metrics for drug-target binding affinity prediction. Our analysis indicates that although many methods report strong performance on standard benchmarks, their effectiveness is often influenced by dataset bias and limited evaluation settings. Furthermore, most methods exhibit reduced performance in cold-start scenarios, highlighting challenges in generalization. We identify several limitations of current approaches, including dataset imbalance, the lack of standardized evaluation, limited real-world applicability, and challenges in cold-start scenarios. We also discuss future research directions, including better dataset design, more robust evaluation methods, improved handling of cold-start problems, and the integration of multimodal representations.
In the early stages of drug discovery, predicting drug-target affinity is a crucial task. Due to the vast scale of genomic and chemical spaces, traditional biological methods are time-consuming, labor-intensive, and resource-demanding. As a result, machine learning-based computational methods have emerged to narrow down the pool of drug candidates. However, machine learning approaches still face several challenges in practical applications, particularly the scarcity of labeled samples and poor model generalization capability. To address these issues, this paper proposes a novel drug-target affinity prediction model, termed MetaBayes-DTA, based on an uncertainty-aware meta-learning framework. The model integrates the few-shot rapid adaptation capability of meta-learning with an uncertainty quantification mechanism to enhance prediction accuracy and reliability. MetaBayes-DTA is evaluated on two benchmark datasets, DAVIS and KIBA. Experimental results demonstrate that the proposed model outperforms existing methods.
Naihan Shi, Yanpeng Zhao, Wanying Li et al.· 2026 IEEE 27th China Confere...· 0 citations
Accurately predicting binding affinities between drugs and targets is crucial for drug discovery but remains challenging due to the complexity of modeling interactions between small drug and large targets. This research presents Dual modality feature fused-drug target affinity (DMFF-DTA), a model for drug-target affinity anticipation using dual-modality neural networks that considers both the sequence and graph structure of medicines and proteins. To facilitate more exact and efficient drug-target interaction modeling, the model incorporates a binding site-focused graph generation method for extracting binding information. Experimental results show that DMFF-DTA is far more effective than current state-of-the-art approaches. By outperforming state-of-the-art approaches by more than 8%, the model demonstrates remarkable generalizability to hitherto unexplored medicines and targets. The model's biological relevance is confirmed by the model interpretability analysis. This paper presents a reliable and understandable method for improving computational drug discovery by integrating multi-view protein and drug properties.
Ghazala Sultan, J. Vincent, Ratna Sahaya et al.· International Conference Com...· 0 citations
This study shows that incorporating synthetic molecular dynamics data improves deep learning models for protein–ligand binding affinity prediction beyond static experimental structures, and highlights that dynamic synthetic datasets can enable deep learning models to outperform conventional methods such as MM-PBSA while remaining computationally efficient.
P. Agrawal, Prathit Chatterjee, U. Priyakumar· Journal of Cheminformatics· 0 citations
Machine learning applications in preclinical drug development have been focused on automated covariate selection in pharmacometric modeling and high-throughput screening processes early in drug discovery. While inherent drug property prediction has made significant improvements in the past decade, fusing early target-based drug discovery methods to preclinical stage pharmacokinetic (PK) property predictions has been limited. This scoping review investigates the current state of PK property prediction of small molecules in drug discovery using machine learning methods and a combination of machine learning and mechanistic models. We identified major obstacles hindering the development of superior prediction models for small molecule behavior in biological systems. These encompass data accessibility, quantity, and quality, architectural constraints such as poor interpretability and model inherent assumptions, and the lack of robust evaluation and uncertainty assessment methods. To mitigate data-related constraints, we advocate for the use of collaborative federated learning frameworks. Furthermore, we propose leveraging the pattern recognition capabilities of deep learning models in conjunction with the biological interpretability provided by mechanistic approaches to strike an optimal balance between accuracy and biological explainability guided by the intended application of the prediction model. Addressing these limitations will advance reliable modeling pipelines and enable effective extrapolation to novel chemical space, additional species, and emerging drug development scenarios.
Lucille Tomin, Vida Bodaghi-Namileh, D. Schwartz et al.· Journal of Chemical Informat...· 0 citations
ABSTRACT Mutation‐induced drug resistance challenges both pandemic surveillance and drug discovery. While experimental assays are resource‐intensive, current computational predictions remain limited by the scarcity of 3D mutant protein structures. We present DeepMutDTA, a structure‐independent model pre‐trained on 1.5 million data points to predict drug‐target affinity and uncover underlying interaction mechanisms. However, like other sequence‐based approaches, it often falls short in predicting mutant affinities due to the overwhelming sequence similarity between wild‐type (WT) and mutant (MT) targets. To bridge this gap, we introduce SimSiam‐MuTF, a novel fine‐tuning framework to enhance the detection of resistance variants by explicitly aligning latent embedding distances with the corresponding shifts in binding affinity between WT and MT targets. Compared to representative baselines, our model exhibits remarkable robustness across varied sequence identities and unseen data splits, yielding average performance gains of 2.47% (PCC) and 5.10% (SCC) in regression tasks, alongside 4.00% (AUC) and 4.17% (AUPR) in classification tasks. Applications to SARS‐CoV‐2, HIV‐1, and cancer‐related targets highlight its generalization potential and utility in informing therapeutic strategies against drug resistance. Collectively, this robust computational pipeline and fine‐tuning framework deepen our understanding of mutation‐induced resistance and may serve as a powerful platform to accelerate drug discovery against mutant targets.
Xiaowen Hu, Pan Zhang, Shangqian Wu et al.· Advancement of science· 0 citations