Skip to content
Review

AI-Driven Drug–Target Interaction Prediction: From Data Representation to Model Design

Aug 2026 · Journal of Chemical Information and Modeling · 0 citations · 151 references

Abstract

Drug–target interaction (DTI) prediction is central to drug discovery, target identification, and drug repurposing. With the rapid growth of biomedical data and advances in artificial intelligence (AI), DTI prediction has shifted from docking, similarity-based inference, and hand-crafted features toward data-driven representation learning and interaction modeling. This review examines AI-driven DTI prediction methodologies, covering binding theories, task formulations, data representation, model design, translational applications, and unresolved challenges. First, we introduce classical molecular binding theories, including the lock-and-key model, induced fit, and conformational selection, highlighting the transition from static matching to dynamic interaction. Second, we summarize major DTI task settings, including binary interaction classification, binding affinity regression, and multitask prediction with uncertainty assessment. We then discuss data resources and multimodal representation approaches for drugs, target proteins, interaction labels, and auxiliary biomedical data, including molecular sequences, graph structures, 3D conformations, physicochemical properties, biological perturbation profiles, protein sequences and structures, and biomedical knowledge networks. Representative approaches are compared across orthogonal methodological dimensions, including input representation, encoder architecture, interaction-modeling mechanism, representation learning and pretraining, learning objective, prediction output, data acquisition or optimization strategy, and generalization setting. Finally, we outline DTI applications in disease target mining, compound virtual screening, affinity and selectivity optimization, complex structure prediction, and industrial drug discovery pipelines and discuss key challenges such as data distribution shifts, dynamic protein conformational variability, and insufficient experimental validation. This survey aims to provide a systematic reference for future algorithm design, mechanism exploration, and real-world drug discovery applications.

View source