Skip to content
Open access

SAPDTA: A Novel Protein Segment Capture Strategy for Drug-Target Affinity Prediction

Jul 2026 · Sensors and AI · 0 citations · 45 references

TL;DR

A novel protein segment capture strategy for drug-target affinity prediction (SAPDTA), which is designed to extract local protein features through a local block capture approach, enabling more flexible extraction of protein structure information at different levels.

Abstract

Predicting drug-target affinity (DTA) is becoming increasingly vital in the field of drug discovery. Currently, many methods focus solely on the overall encoding of proteins, overlooking the abundant information contained within protein peptides. Therefore, this paper proposes a novel protein segment capture strategy for drug-target affinity prediction (SAPDTA), which is designed to extract local protein features through a local block capture approach. This strategy supports adaptive segmentation of amino acid chains, enabling more flexible extraction of protein structure information at different levels. A hybrid dual-network bilinear interaction module is proposed to address the challenge of protein feature extraction at various scales. Moreover, bilinear interaction blocks are employed to combine and process the chemical properties of drugs with the biological characteristics of their targets. SAPDTA’s performance is assessed using two publicly accessible DTA datasets (Davis and KIBA). According to the experimental results, SAPDTA demonstrates competitive performance compared to existing models across all evaluation metrics. Furthermore, visualization results on the ToxCast dataset highlight the model’s sensitivity to complex drug structures, revealing its capability to understand underlying structure-function relationships.

Read PDF

Similar papers

Jul 2026

PSDTA: An Approach to Drug-Target Binding Affinity Prediction by Integrating Physicochemical and Structural Information to Reduce Feature Redundancy.

This work proposes a DTA prediction method, PSDTA, which integrates physicochemical properties into the initial feature representations and explicitly incorporates structural information on amino acids, thereby avoiding the risk of information leakage caused by directly using coordinates as features and enhancing the model's generalization capability.

Shuang Wang, Mao Li, Peifu Han et al. · 0 citations
Aug 2026

MMU-DPI: Enhancing Generalization in Drug–Protein Interaction Prediction through Multimodal Learning and a Label Mix Strategy

Accurate prediction of drug–protein interactions (DPIs) is crucial for accelerating the drug discovery process. However, the scarcity of experimentally validated interactions can limit the learning of transferable interaction patterns, particularly for previously unseen drugs and proteins. To address this fundamental challenge, we propose the MMU-DPI framework. A key component of this framework is a Label Mix strategy tailored to multimodal DPI prediction, which performs interpolation only in the label space while keeping the input modalities unchanged. This strategy provides stochastic soft-target regularization and improves generalization performance under reduced-data and independent external Cold-both evaluation settings. To effectively process and utilize multimodal data, MMU-DPI adopts a multimodal dual-branch architecture. The first branch uses a Message Passing Neural Network (MPNN) to extract structured representations from drug molecular graphs. It also uses a Convolutional Neural Network (CNN) to capture key biological and functional features from amino acid sequences. The second branch constructs a heterogeneous interaction graph and uses a Graph Attention Network (GAT) to learn deep contextual relationships between drugs and proteins. A learnable global fusion weight combines complementary branch logits to generate the final prediction for each drug–protein pair. Experimental results on multiple benchmark data sets demonstrate that MMU-DPI outperforms several state-of-the-art DPI prediction methods. Case studies further support the ability of MMU-DPI to identify potential DPIs. These results indicate that MMU-DPI can serve as a useful computational tool for drug discovery.

Jiahao Wei, Tie Shen · 0 citations
#protein folding Open access Aug 2026

DyAb: sequence-based antibody design and property prediction in a low-data regime.

Protein therapeutic design and property prediction are frequently hampered by data scarcity. Here we propose a model, DyAb, that addresses these issues by leveraging a pair-wise representation to predict differences in binding affinity, rather than absolute values. DyAb is built on top of a pre-trained protein language model and achieves a Spearman rank correlation of up to 0.85 on binding affinity prediction across monoclonal antibodies targeting three different antigens (EGFR, IL-6, and an internal target), given as few as 100 training data. We employ DyAb in two design contexts: as a ranking model to score combinations of known mutations, and combined with a genetic algorithm to generate new sequences. Our method consistently generates antibody variants with high binding rates, including designs that improve on the binding affinity of the lead molecule by more than ten-fold. DyAb represents a powerful tool for optimizing antibody binding affinity in low data regimes common in early-stage drug development.

J. Lin, Jennifer L. Hofmann, Andrew Leaver-Fay et al. · 0 citations
Jun 2025

READ: A Retrieval-Alignment Diffusion Framework for Structure-based Drug Design.

Structure-based drug design (SBDD) models are central to modern pharmaceutical research, enabling the rational exploration of protein-ligand interactions at atomic resolution. However, most existing approaches frame molecular generation as an isolated optimization or a one-to-one matching task, overlooking the shared binding patterns and intrinsic similarities among protein-ligand complexes. This fragmented perspective constrains their ability to capture the fundamental principles governing molecular recognition and binding specificity. Moreover, the limited availability of high-quality experimental data further hampers model generalization and real-world applicability. To address these challenges, we present READ, a retrieval-alignment molecular generation framework that conditions the generative process on small molecules targeting homologous proteins. Retrieved ligands are aligned with a diffusion model across multiple representational spaces and integrated as conditional guidance throughout successive stages of generation. Under a standardized docking-based evaluation protocol, READ achieves consistently strong performance against state-of-the-art SBDD methods. More importantly, it introduces a retrieval-alignment paradigm for structure-based molecular generation, offering a practical framework for early-stage computational hit generation while leaving prospective experimental validation as future work.

Dong Xu, Zhangfan Yang, Junchuang Cai et al. · 1 citation
Jul 2026

A Scalable Structure-Aware Multimodal Architecture for Accurate Drug-Target Affinity Prediction.

Accurate prediction of drug-target binding affinity (DTA) is a key task in virtual screening. However, current computational methods face a key challenge: sequence-based approaches often fail to capture critical spatial information, while structure-based models rely on computationally expensive 3D coordinates, which restrict their scalability. To address this issue, we propose StructuraDTA, a novel multimodal framework that adopts an implicit structure modeling strategy. Instead of using static protein folding data, our method encodes drug molecular graphs via Graph Isomorphism Networks (GINs) to capture fine-grained topological features. Meanwhile, we optimize protein representations by integrating probabilistic structural priors into a pretrained language model, which effectively simulates thermodynamic conformational flexibility without relying on explicit 3D structural data. A bidirectional cross-attention mechanism is then used to dynamically align these heterogeneous feature modalities. Comprehensive evaluations on the Davis and KIBA benchmark datasets show that StructuraDTA stably outperforms state of-the-art comparison methods. Importantly, the model exhibits strong robustness in cold-start scenarios, and can accurately predict binding affinities for previously unseen drugs and targets. By retaining the predictive performance of structure based models while maintaining the high inference efficiency of sequence-based methods, we provide an accurate and scalable solution to accelerate genome-scale drug discovery research.

Junlin Xu, Ye Yuan, Menglong Hu et al. · 0 citations
Conference Open access 2026

Prediction of Drug Target Interactions Based on Artificial Intelligence

Drug-target interaction (DTI) prediction is an efficient pre-screening method that uses algorithmic models to assess the binding potential of drug molecules to protein targets. Current research is accelerating towards the integration of heterogeneous graph neural networks, protein language models, and generative artificial intelligence. This review systematically summarizes the latest developments in these technologies, pointing out the problems currently being addressed in research such as data sparsity and cold start, as well as the manifestations of general machine learning challenges such as recommendation systems and noise learning in the biomedical field; Interpret representative models such as Dual Heterogeneous Graph Transformer for Drug-Target Interaction (DHGT-DTI) and Graph Positional encoding and Sequence features for Drug-Target Interaction (GPS-DTI). The combination of dual perspective learning, equivariant graph convolution, and attention mechanism enhances the understanding and reasoning ability of these models in complex biological networks. The generative artificial intelligence diffusion model has opened up a path for developing new drugs through structured and data enhanced approaches. Research has shown that important issues related to computational performance, interpretability, and data compatibility still need to be addressed in existing models. Building a high-performance, multifunctional pre trained model for large-scale biomolecules should be a key direction for development.

Qizhong Yang · 0 citations