Skip to content
Open access

Development and validation of an attention-based cGAS-specific deep learning scoring function for structure-based virtual screening

Jul 2026 · Discover Chemistry · Vol 3 · 0 citations · 34 references

TL;DR

DeepCGASPred is a cyclic GMP-AMP synthase (cGAS)-specific deep learning scoring function that integrates three-dimensional convolutional neural networks with multi-head attention mechanisms and composite structural descriptors, including Structural Protein–Ligand Interaction Fingerprints (SPLIF), hydrogen bond features, and extended connectivity fingerprints (ECFP).

Abstract

Structure-based virtual screening (SBVS) is a cornerstone of modern drug discovery pipelines, yet conventional scoring functions often lack the resolution to model complex protein–ligand interactions accurately. To overcome these limitations, we developed DeepCGASPred, a cyclic GMP-AMP synthase (cGAS)-specific deep learning scoring function that integrates three-dimensional (3D) convolutional neural networks (CNNs) with multi-head attention mechanisms and composite structural descriptors, including Structural Protein–Ligand Interaction Fingerprints (SPLIF), hydrogen bond features, and extended connectivity fingerprints (ECFP). This integrative approach enables the model to capture spatial, physicochemical, and topological interaction patterns while prioritizing informative regions for accurate classification of active and inactive compounds. DeepCGASPred was trained using a chemically diverse dataset and rigorously validated through systematic hyperparameter optimization and multiple independent runs. Our results demonstrate that the combined feature set SPLIF+Hbonds+ECFP consistently outperforms all tested feature configurations, achieving a precision-recall area under the curve (PR-AUC) of 0.94–0.97, a median precision of approximately 0.99, a median recall of approximately 0.88, and a median F1 score above 0.92 across ten independent runs. Comparative evaluation reveals that DeepCGASPred surpasses established scoring functions such as SMINA, RF-Score, SCORCH, and CNN-Score on this cGAS-specific dataset, particularly in identifying actives under challenging test conditions. A perfect normalized enrichment factor at 1% (NEF1% = 1.00) confirms strong early enrichment performance. Optimal performance was achieved when the attention mechanism was inserted after the first CNN layer, reinforcing its role in enhancing generalization. DeepCGASPred offers a robust, interpretable framework for target-specific SBVS of cGAS inhibitors; the approach may also inform the development of similar tools for other biological targets.

Read PDF

Similar papers

Open access Jul 2026

3Br-MGD: few-shot toxicity prediction with a three-branch deep encoder and meta-learning framework.

Predicting the toxicity of pharmaceutical compounds remains a major challenge in drug discovery. Early and accurate toxicity assessment is essential for eliminating harmful candidates before costly preclinical and clinical testing, thereby improving patient safety, reducing development costs, and accelerating the drug development process. Despite advances in computational toxicology, existing methods often struggle to capture complex molecular characteristics and maintain robust performance under limited-data conditions. To address these challenges, we propose 3Br-MGD, a novel three-branch framework that integrates deep learning and meta-learning for molecular toxicity prediction. The architecture combines complementary molecular representations: FingerprintMLP encodes Morgan fingerprint descriptors, Graph Convolutional Networks (GCNs) capture structural information from molecular graphs, and one-dimensional Deep Convolutional Neural Networks (1D-CNNs) extract sequential features from SMILES strings. These embeddings are integrated within a Prototypical Network-based few-shot learning framework, enabling rapid adaptation to new prediction tasks with limited labeled samples and improving generalization in low-resource settings. Experimental results on benchmark toxicity datasets demonstrate that 3Br-MGD consistently outperforms conventional baselines in predictive accuracy, robustness, and generalization. Furthermore, the integration of heterogeneous molecular encoders reduces dependence on large training datasets while enhancing interpretability through the exploitation of complementary chemical information from multiple molecular views.

Nguyen Thi Phuong Thao, Bui Thanh Hung · 0 citations
Aug 2026

Dual-Attention Multimodal Framework for Molecular Property Prediction

A novel Dual-Attention Multimodal framework for Graphs and Sequence-based representations, so-called DAM-GS, which provides a promising solution for molecular property prediction with broad applications in drug discovery and computational molecular science.

Bay Van Nguyen, Vinh Truong, Ha Duong Thi Hong et al. · 0 citations
Jul 2026

A synergistic deep learning and machine learning framework for screening heterocyclic compounds against ALDH1A1.

An integrated computer-aided drug design (CADD) and artificial intelligence (AI) framework to systematically identify selective ALDH1A1 inhibitors from a heterocyclic compound library is developed and suggests that LDN-27219 exhibits favorable binding characteristics and represents a promising lead candidate for subsequent experimental validation.

Shu-Chi Cho, Yi-Wen Wang, Chien-An Chu et al. · 0 citations
Open access Aug 2026

CLDN18.2 antibody design with protein language models: A deep learning optimization framework

CLDN18.2 is a promising tumor-specific antigen; however, the development of therapeutic antibodies against it is challenged by the need for simultaneous optimization of affinity and developability. To address this, we present cdrGPT, a deep learning framework based on GPT-2 for de novo generation of complementarity-determining region H3 (CDRH3) sequences. Our approach integrates pre-training on the Observed Antibody Space (OAS) database with structural templating derived from the known antibody zolbetuximab. Generated sequences were iteratively refined through rejection sampling and fine-tuned against a multi-parameter objective function encompassing predicted affinity and MHC class II binding risk. From an initial set of 50,000 sequences, this screening pipeline yielded 313 high-confidence candidates. Subsequent analysis using evolutionary scale modeling 2 (ESM2) embeddings, principal component analysis (PCA), and clustering revealed three structurally distinct clusters, with intra-cluster cosine similarities exceeding 0.99. Validation of seven representative sequences from the dominant cluster using AlphaFold3 confirmed high structural fidelity to the zolbetuximab template, demonstrating a root mean square deviation (RMSD) of 1.331 Å for the CDRH3 loop and positional deviations of less than 0.4 Å for key paratope residues. These results indicate that the designed variants preserve the core binding mode of the parent antibody. This study establishes a feasible pipeline for integrating AI-generated CDRH3 loops into functional antibody scaffolds, providing a foundation for the accelerated development of therapeutics targeting CLDN18.2 and other clinically relevant antigens.

Tao Qu, Lingyan Yuan, Weiran Cui et al. · 0 citations
Jul 2026

A Scalable Structure-Aware Multimodal Architecture for Accurate Drug-Target Affinity Prediction.

Accurate prediction of drug-target binding affinity (DTA) is a key task in virtual screening. However, current computational methods face a key challenge: sequence-based approaches often fail to capture critical spatial information, while structure-based models rely on computationally expensive 3D coordinates, which restrict their scalability. To address this issue, we propose StructuraDTA, a novel multimodal framework that adopts an implicit structure modeling strategy. Instead of using static protein folding data, our method encodes drug molecular graphs via Graph Isomorphism Networks (GINs) to capture fine-grained topological features. Meanwhile, we optimize protein representations by integrating probabilistic structural priors into a pretrained language model, which effectively simulates thermodynamic conformational flexibility without relying on explicit 3D structural data. A bidirectional cross-attention mechanism is then used to dynamically align these heterogeneous feature modalities. Comprehensive evaluations on the Davis and KIBA benchmark datasets show that StructuraDTA stably outperforms state of-the-art comparison methods. Importantly, the model exhibits strong robustness in cold-start scenarios, and can accurately predict binding affinities for previously unseen drugs and targets. By retaining the predictive performance of structure based models while maintaining the high inference efficiency of sequence-based methods, we provide an accurate and scalable solution to accelerate genome-scale drug discovery research.

Junlin Xu, Ye Yuan, Menglong Hu et al. · 0 citations