Skip to content

ProphDR: An Interpretable Deep Learning Model for Predicting Cancer Drug Response via Multi-Omics and Cross-Attention Mechanisms.

Jul 2026 · Journal of Chemical Information and Modeling · Vol 66, pp. 7874-7888 · 0 citations · 47 references
Medicine

TL;DR

ProphDR is an interpretable deep learning framework that integrates multiomics data and drug structural information using a hierarchical attention mechanism, and generates biologically interpretable attention maps that highlight key pharmacophores and resistance-related genes consistent with established mechanisms in NSCLC and BRCA.

Abstract

Predicting cancer drug responses (CDRs) accurately remains a significant challenge due to the complexity of tumor biology and the limitations of existing "black-box" machine learning models. To address this, we propose ProphDR, an interpretable deep learning framework that integrates multiomics data and drug structural information using a hierarchical attention mechanism. ProphDR incorporates a Criss-Cross Gene-level Multiomics Integration (CGMI) module to capture gene-level features and a cross-attention (CA) module to model drug-gene interactions. Evaluated on datasets from GDSC and CCLE, ProphDR achieves state-of-the-art performance in predicting ln(IC50) values (PCC = 0.938, RMSE = 0.978) and classifying drug sensitivity (AUC = 0.981). It also demonstrates strong generalizability in cold-start scenarios involving unseen drugs or cell lines. Crucially, ProphDR generates biologically interpretable attention maps that highlight key pharmacophores and resistance-related genes such as ERBB2 (HER2), consistent with established mechanisms in NSCLC and BRCA. These insights bridge genomic features with phenotypic outcomes, offering valuable guidance for target prioritization and drug repurposing. ProphDR represents a robust and explainable AI tool for advancing precision oncology.

View source

Similar papers

Open access Jul 2026

Separate XAI: Independent Training Framework for Cancer Drug Sensitivity Prediction Using GDSC and CCLE with Explainable AI-Driven Drug Repositioning

Background: The high costs, long development timelines, and low clinical success rates in oncology highlight an urgent need for reliable computational strategies for drug repositioning. Current machine learning approaches often integrate heterogeneous pharmacogenomic datasets, which may lose biological specificity and limit model interpretability. Methods: In this study, we propose Separate XAI, an explainable artificial intelligence framework that retains dataset-specific biological features by adopting separate preprocessing and training pipelines for the Genomics of Drug Sensitivity in Cancer (GDSC) and Cancer Cell Line Encyclopedia (CCLE) datasets. Different deep learning architectures such as Deep Neural Networks (DNNs), Convolutional Neural Networks (CNNs), and Recurrent Neural Networks (RNNs) were used to predict the drug response in the cancer cell lines. We also used SHapley Additive exPlanations (SHAP) to improve interpretability and identify biologically relevant features. Results: The developed framework showed good predictions with 94.49% accuracy in the CCLE dataset and a mean squared error of 0.0725 in the GDSC dataset. Explainability analysis identified important biomarkers and signaling pathways such as TP53 and KRAS, providing mechanistic insights into drug sensitivity and therapeutic response. Conclusions: The distinct XAI presented here offers an interpretable, biologically grounded framework for cancer drug repositioning by integrating dataset-specific modeling and explainable artificial intelligence. However, integration-based approaches often suffer from confounding effects of experimental and biological heterogeneity, but the proposed framework explicitly preserves dataset-specific characteristics, which potentially could lead to more robust predictions and higher interpretability for precision oncology and translational cancer research.

Heba M. Nagy, F. Maghraby, Osama M. Badawy et al. · 0 citations
Conference Jul 2026

Explainable Multi-Omic Machine Learning Framework for Predicting Drug Response in Breast Cancer

Accurate prediction of drug sensitivity in cancer cell lines is vital for precision oncology and patient-specific therapies. However, many computational approaches fail to integrate multi-modal biological and chemical features and often struggle with high-dimensional, imbalanced pharmacogenomic data, limiting predictive accuracy and interpretability. To address these challenges, we developed a machine learning framework that integrates pharmacogenomic profiles-including mutation status, copy number alterations, and microsatellite instabil-ity-with molecular fingerprints and descriptors of 85 anticancer drugs, generated using PaDEL from SMILES strings. Data from 40 breast cancer cell lines in the Genomics of Drug Sensitivity in Cancer (GDSC) dataset were employed. A threestage feature selection strategy combining Boruta, mRMR, and XGBoost was applied to reduce drug feature dimensionality while retaining 130 cell line features. Multiple models were trained, and LightGBM, optimized with grid search, class weighting, and 3-fold cross-validation, demonstrated superior performance in handling severe class imbalance (233 sensitive vs. 3167 resistant samples). LightGBM achieved training AUROC $=0.9455$, AUPRC $\boldsymbol{=} \mathbf{0. 5 1 4 8}$, Accuracy $\boldsymbol{=} \mathbf{0. 8 4 1 5}$, F1-score = 0.4481, Recall = 0.9409, and MCC = 0.4732, underscoring its suitability for sparse biomedical datasets. Model interpretation with SHapley Additive exPlanations (SHAP) highlighted BRCA-related features, identifying cnaBRCA25 (not mutated) as a resistance marker and cnaBRCA47 (mutated) as a context-dependent biomarker, consistent with their roles in DNA repair pathways. Overall, this framework demonstrates the value of multi-modal integration and interpretable machine learning in pharmacogenomics. While results are promising, validation on larger and independent cohorts is essential to establish clinical relevance.

D. Kumari, Aiman, Sakshi Singh et al. · 0 citations
Review Open access Jul 2026

DR.DEGMON: self-explainable deep neural network for drug-induced cell viability prediction incorporating differentially expressed genes and gene ontology

Accurate prediction of cancer drug responses is essential for advancing cancer treatment strategies and drug development. With the increasing availability of large-scale pharmacogenomic datasets, many deep learning models have been proposed to predict cancer drug responses. However, many existing models lack the capacity to offer critical biomedical insights, such as providing interpretability regarding the potential mechanism of action. We propose DR.DEGMON (Drug Response prediction using Differentially Expressed Genes with Multi-layer perceptron integrating gene Ontology Network), a self-explainable deep neural network designed to predict the viability of pan-cancer cell lines in response to drug treatments by utilizing differentially expressed genes. DR.DEGMON leverages prior biological knowledge by incorporating Gene Ontology (GO) into the hierarchical structure of a multi-layer perceptron. The architecture of DR.DEGMON highlights key genes and GO terms that contribute to drug responses through layer-wise relevance propagation (LRP), suggesting potential biological pathways associated with specific drugs. DR.DEGMON achieved a Pearson correlation coefficient of 0.8568 for cell viability prediction, outperforming all baseline models. The model also showed robust generalization performance on external datasets, including GDSC, PRISM, and CCLE. In addition, we employed layer-wise relevance propagation (LRP) to obtain relevance scores for input genes and nodes representing GO terms. DR.DEGMON shows high performance in predicting drug responses and provides interpretable results. The integration of GO and LRP enabled the model to suggest the underlying biological processes involved in drug responses, making it a valuable tool for predicting outcomes and discovering new biomedical knowledge in cancer pharmacogenomics. This approach offers both practical utility in drug development and a method for improving the understanding of cancer biology.

Wootaek Lim, Jitae Kim, Songhyeon Kim et al. · 0 citations
Open access Aug 2026

ZSCAN-DDIE: An interpretable zero-shot learning method for the prediction of drug-drug interaction events using biomedical text

Unexpected drug–drug interaction events (DDIEs) pose substantial clinical risks, yet many remain unannotated due to data scarcity and the rapid emergence of novel drug combinations. Conventional deep learning approaches struggle to generalize to these unseen interaction types and often lack interpretability under severe class imbalance. To address these challenges, we propose ZSCAN-DDIE, an interpretable zero-shot learning framework for DDIE prediction. The model integrates a biomedical pre-trained language model with an attention-based graph convolutional network (AGCN) to encode DDIE textual semantics and drug molecular structures, respectively. A cross-attention network (CAN) is introduced to align molecular substructures with pharmacological semantic components, enabling fine-grained cross-modal reasoning and improving interpretability at the substructure level. To mitigate modality bias and long-tailed distribution effects, we design a bimodal dynamic alignment (BDA) loss that combines hyperspherical embedding regularization with a stage-adaptive loss-switching mechanism. Experimental results under both conventional and generalized zero-shot settings demonstrate that ZSCAN-DDIE consistently outperforms state-of-the-art baselines across multiple evaluation metrics. The proposed framework not only enhances prediction accuracy for unseen DDIE categories but also provides biologically meaningful insights into molecular interaction mechanisms, offering a robust and clinically relevant solution for pharmacovigilance and drug safety assessment The source code and data are available at https://github.com/GSX-0429/ZSCAN-DDIE.

Shaoxi Gao, Zhanpeng Gan, Fangfang Han et al. · 0 citations
Open access Aug 2026

Drug sensitivity prediction across cancer types using graph isomorphism networks and biological pathway features: A dual-branch deep learning approach

Drug sensitivity prediction is an important issue within the precision medicine field. IC50, which is the molar drug dose needed to decrease the viability of cells by half compared to the drug-free control, is the main pharmacodynamics parameter used for drug sensitivity analysis in large-scale pharmacogenomics screenings. Computational estimation of IC50s based on molecular and genomic factors significantly reduces costs associated with experiments for measuring cell viability and allows for accelerating the process of drug discovery. Traditional methods of IC50 calculation do not allow integrating the three-dimensional chemical structure of drugs and the biological context of particular cell lines, resulting in suboptimal model performance when using different pharmacogenomics data sources. In this work, we propose an innovative dual-branch approach based on Graph Isomorphism Network (GIN) drug representations coupled with a Multilayer Perceptron (MLP) for 50-dimensional ssGSEA pathway activities calculated from CCLE gene expression. After training on cell-line-drug pair combinations from the Genomics of Drug Sensitivity in Cancer 2 (GDSC2) dataset across various cancers, the proposed GIN+Pathway MLP model attains an R2 of 0.8553 and a Pearson Correlation Coefficient (PCC) of 0.9249 on the testing split of the same dataset. In a variant ablation study of six variants, we find that eliminating the pathway MLP component lowers the R2 value by more than 0.15, thus proving the importance of biological features in the two-branch model. The performance of our proposed model exceeds benchmark scores for models such as GraphDRP (PCC = 0.870, R2 = 0.756) and DeepCDR (PCC = 0.847, R2 = 0.720) when tested on the same GDSC2 dataset.

Shuang Li, Quanzhong Yang, Feifei Shen et al. · 0 citations
Open access Jul 2026

Essentiality-driven prediction of anticancer drug responses in preclinical and clinical contexts

Summary Precision oncology relies on tumor molecular profiles to predict drug responses. Instead of using conventional molecular features directly, we construct predictive signatures based on gene essentiality. Here, we present DrGee, an essentiality-centered platform that infers drug sensitivity solely from gene expression profiles. The built-in DeepEEAA model integrates gene expression, gene essentiality, drug-protein affinity, and drug-gene associations to quantitatively predict IC50 values. DeepEEAA achieved competitive predictive performance on independent cell line datasets (R2 = 0.764; MSE = 0.9345), outperforming recent benchmark deep learning methods. DrGee prioritized four candidate drugs for the 95-D lung cancer cell line, among which BI-97C1 and trimetrexate were validated by in vitro assays and mouse xenograft experiments. Robust predictive performance was further confirmed in OVCAR8 ovarian cancer cells. In TCGA cohorts, essentiality-driven predictions stratified patients with significantly different overall survival outcomes (AUC-PR = 0.825), highlighting the translational potential of DrGee.

Hongtu Cui, Xiaohui Du, Hai-Xia Guo et al. · 0 citations