Accurate prediction of drug-target affinity (DTA) is essential for accelerating drug discovery. Although pretrained protein language models have achieved significant progress, existing methods predominantly focus on bottom-up sequence patterns and lack explicit constraints from high-level biological functions. We propose GoMA-DTA, a framework integrating gene ontology (GO) functional annotations with protein semantic features. GoMA-DTA introduces a channelwise gating mechanism that uses functional semantics as anchors to dynamically recalibrate ESM-2embeddings, achieving adaptive semantic filtering. For drugs, the model integrates Molformer-based semantic and TransConv-derived structural features. These dual-modality drug representations interact with calibrated protein features through a parallel synergistic architecture of cross-attention and Mamba modules, ensuring precise cross-modal alignment and efficient long-range dependency modeling. Evaluations on PDBBind, BindingDB, and ChEMBL benchmarks demonstrate that GoMA-DTA significantly outperforms state-of-the-art models across various evaluation scenarios. Its superior screening power is further validated on CASF-2016. Moreover, virtual screening of 200 000compounds against the SARS-CoV-2Spike protein, supported by experimental evidence (ZINC2111387), underscores its practical utility as a robust and biologically reliable tool. The datasets and codes are publicly available at https://github.com/xa-123955/GoMA-DTA.
An Xiong, Zheyu Zhou, Yazi Li et al.· IEEE Transactions on Neural...· 0 citations
MultiGeo is a DTA prediction framework that explicitly leverages multiple protein conformations rather than a single snapshot, and introduces a disagreement-aware gating mechanism that adaptively fuses this ensemble representation with the dominant structure only when the additional conformers provide complementary information.
Ruida Zeng, Cheng Guo, Yajie Meng et al.· 0 citations
Accurate prediction of drug-target binding affinity (DTA) is a key task in virtual screening. However, current computational methods face a key challenge: sequence-based approaches often fail to capture critical spatial information, while structure-based models rely on computationally expensive 3D coordinates, which restrict their scalability. To address this issue, we propose StructuraDTA, a novel multimodal framework that adopts an implicit structure modeling strategy. Instead of using static protein folding data, our method encodes drug molecular graphs via Graph Isomorphism Networks (GINs) to capture fine-grained topological features. Meanwhile, we optimize protein representations by integrating probabilistic structural priors into a pretrained language model, which effectively simulates thermodynamic conformational flexibility without relying on explicit 3D structural data. A bidirectional cross-attention mechanism is then used to dynamically align these heterogeneous feature modalities. Comprehensive evaluations on the Davis and KIBA benchmark datasets show that StructuraDTA stably outperforms state of-the-art comparison methods. Importantly, the model exhibits strong robustness in cold-start scenarios, and can accurately predict binding affinities for previously unseen drugs and targets. By retaining the predictive performance of structure based models while maintaining the high inference efficiency of sequence-based methods, we provide an accurate and scalable solution to accelerate genome-scale drug discovery research.
Junlin Xu, Ye Yuan, Menglong Hu et al.· IEEE journal of biomedical a...· 0 citations
MolLingua, a token-centric, dual-modal framework designed for native molecular understanding, uses a dual-branch Residual Vector Quantization engine to discretize heterogeneous, high-dimensional spatial 2D and 3D features into compact code sequences rather than relying solely on continuous projections.
Haoyang Liu, Xikang Feng, Fei Guo et al.· IEEE journal of biomedical a...· 0 citations