Jul 2026· Journal of Chemical Information and Modeling· 0 citations· 45 references
Medicine
TL;DR
This work proposes a DTA prediction method, PSDTA, which integrates physicochemical properties into the initial feature representations and explicitly incorporates structural information on amino acids, thereby avoiding the risk of information leakage caused by directly using coordinates as features and enhancing the model's generalization capability.
Abstract
Accurate prediction of drug-target binding affinity (DTA) is critical for repurposing. Although deep learning has been widely applied in this field, existing methods still face challenges, including inadequate integration of physicochemical properties and structural information, which leads to high feature redundancy and limited generalization performance. To address these challenges, we propose a DTA prediction method, PSDTA (where P represents physicochemical properties and S represents structural information), which integrates physicochemical properties into the initial feature representations. Unlike conventional approaches, PSDTA explicitly incorporates structural information on amino acids, thereby avoiding the risk of information leakage caused by directly using coordinates as features and enhancing the model's generalization capability. Furthermore, two complementary channels are designed to identify binding-relevant residues at both residue and group levels, thereby reducing feature redundancy. To evaluate its effectiveness, we conducted experiments on three benchmark data sets: PDBBind v2016, PDBBind v2020, and Davis. By comparing with several state-of-the-art algorithms, PSDTA achieves the best results in terms of performance. Interpretability analyses indicate that the two channels identify highly consistent and complementary binding-related regions.
A novel protein segment capture strategy for drug-target affinity prediction (SAPDTA), which is designed to extract local protein features through a local block capture approach, enabling more flexible extraction of protein structure information at different levels.
Zihao Fang, Guanqiu Qi, Stanley Tang et al.· Sensors and AI· 0 citations
Accurate prediction of drug-target binding affinity (DTA) is a key task in virtual screening. However, current computational methods face a key challenge: sequence-based approaches often fail to capture critical spatial information, while structure-based models rely on computationally expensive 3D coordinates, which restrict their scalability. To address this issue, we propose StructuraDTA, a novel multimodal framework that adopts an implicit structure modeling strategy. Instead of using static protein folding data, our method encodes drug molecular graphs via Graph Isomorphism Networks (GINs) to capture fine-grained topological features. Meanwhile, we optimize protein representations by integrating probabilistic structural priors into a pretrained language model, which effectively simulates thermodynamic conformational flexibility without relying on explicit 3D structural data. A bidirectional cross-attention mechanism is then used to dynamically align these heterogeneous feature modalities. Comprehensive evaluations on the Davis and KIBA benchmark datasets show that StructuraDTA stably outperforms state of-the-art comparison methods. Importantly, the model exhibits strong robustness in cold-start scenarios, and can accurately predict binding affinities for previously unseen drugs and targets. By retaining the predictive performance of structure based models while maintaining the high inference efficiency of sequence-based methods, we provide an accurate and scalable solution to accelerate genome-scale drug discovery research.
Junlin Xu, Ye Yuan, Menglong Hu et al.· IEEE journal of biomedical a...· 0 citations
The proposed MSIGR-PLA is an integrative framework that integrates local multi-scale interaction features with global protein-ligand representations to improve the accuracy of PLA prediction and consistently outperforms existing methods on four benchmark datasets.
Hangchen Zhang, Haoran Chen, Chang Liu et al.· IEEE journal of biomedical a...· 0 citations
Drug repurposing offers a time-efficient strategy for identifying therapeutics against emerging pathogens such as SARS-CoV-2. In this study, we apply MPS2IT-DTI (Molecule and Protein Sequence to Image Transformer for Drug-Target Interaction), a deep learning framework that represents molecular (SMILES) and protein (FASTA) sequences as images using k-mer frequency encoding, enabling convolutional neural networks to capture spatial compositional patterns associated with biochemical interactions. A curated dataset (BindingDB-FDA) containing 83,165 binding interactions from 1640 FDA-approved ligands and 3270 targets was constructed from BindingDB, with binding scores derived from the KIBA scoring system. An enhanced variant, MPS2IT+MN, incorporating max-norm regularization, was introduced to improve generalization. The model was applied to predict binding affinities between 33 FDA-approved antiviral drugs and six key SARS-CoV-2 non-structural proteins. Results consistently identified five antivirals - MK-5172 (Grazoprevir), Simeprevir, Lopinavir, Etravirine, and Atazanavir - as top-ranked candidates across all targets. Comparative analysis with the MT-DTI model demonstrated competitive and, in several cases, superior ranking performance despite a simpler architecture. Importantly, these predictions are supported by independent experimental and clinical evidence, highlighting the potential of image-based representations as a computationally efficient and biologically meaningful approach for drug-target interaction prediction and drug repurposing.
Jackson G de Souza, Marcelo A. C. Fernandes, Raquel de Melo Barbosa· Computational biology and ch...· 0 citations
Protein therapeutic design and property prediction are frequently hampered by data scarcity. Here we propose a model, DyAb, that addresses these issues by leveraging a pair-wise representation to predict differences in binding affinity, rather than absolute values. DyAb is built on top of a pre-trained protein language model and achieves a Spearman rank correlation of up to 0.85 on binding affinity prediction across monoclonal antibodies targeting three different antigens (EGFR, IL-6, and an internal target), given as few as 100 training data. We employ DyAb in two design contexts: as a ranking model to score combinations of known mutations, and combined with a genetic algorithm to generate new sequences. Our method consistently generates antibody variants with high binding rates, including designs that improve on the binding affinity of the lead molecule by more than ten-fold. DyAb represents a powerful tool for optimizing antibody binding affinity in low data regimes common in early-stage drug development.
J. Lin, Jennifer L. Hofmann, Andrew Leaver-Fay et al.· mAbs· 0 citations