Skip to content
Conference

Explainable Multi-Omic Machine Learning Framework for Predicting Drug Response in Breast Cancer

Jul 2026 · Annual International Computer Software and Applications Conference · pp. 744-753 · 0 citations · 19 references

Abstract

Accurate prediction of drug sensitivity in cancer cell lines is vital for precision oncology and patient-specific therapies. However, many computational approaches fail to integrate multi-modal biological and chemical features and often struggle with high-dimensional, imbalanced pharmacogenomic data, limiting predictive accuracy and interpretability. To address these challenges, we developed a machine learning framework that integrates pharmacogenomic profiles-including mutation status, copy number alterations, and microsatellite instabil-ity-with molecular fingerprints and descriptors of 85 anticancer drugs, generated using PaDEL from SMILES strings. Data from 40 breast cancer cell lines in the Genomics of Drug Sensitivity in Cancer (GDSC) dataset were employed. A threestage feature selection strategy combining Boruta, mRMR, and XGBoost was applied to reduce drug feature dimensionality while retaining 130 cell line features. Multiple models were trained, and LightGBM, optimized with grid search, class weighting, and 3-fold cross-validation, demonstrated superior performance in handling severe class imbalance (233 sensitive vs. 3167 resistant samples). LightGBM achieved training AUROC $=0.9455$, AUPRC $\boldsymbol{=} \mathbf{0. 5 1 4 8}$, Accuracy $\boldsymbol{=} \mathbf{0. 8 4 1 5}$, F1-score = 0.4481, Recall = 0.9409, and MCC = 0.4732, underscoring its suitability for sparse biomedical datasets. Model interpretation with SHapley Additive exPlanations (SHAP) highlighted BRCA-related features, identifying cnaBRCA25 (not mutated) as a resistance marker and cnaBRCA47 (mutated) as a context-dependent biomarker, consistent with their roles in DNA repair pathways. Overall, this framework demonstrates the value of multi-modal integration and interpretable machine learning in pharmacogenomics. While results are promising, validation on larger and independent cohorts is essential to establish clinical relevance.

View source

Similar papers

Open access Aug 2026

Identifying multi-omics biomarkers for ovarian cancer survival estimation

Ovarian cancer is among the deadliest gynecologic malignancies, and its molecular heterogeneity limits accurate prognostic stratification. Although multi-omics approaches have improved predictive modeling, many prioritize predictive performance over biological interpretability, limiting their clinical translation. We developed an interpretable three-stage machine learning framework integrating mRNA, microRNA, DNA methylation, copy number variation, and protein expression data from The Cancer Genome Atlas. Hierarchical feature selection was combined with a weighted ensemble of ElasticNet, ridge regression, support vector regression, XGBoost, and random forest models to estimate overall survival time in patients with ovarian cancer. Multi-omics integration outperformed every single-modality model, achieving a Pearson correlation of 0.752, a concordance index of 0.779, and a mean absolute error of 8.57 months between estimated and observed survival time, compared with 0.48 for the best single modality. The framework identified a 20-biomarker signature dominated by tumor-associated macrophage and complement genes. In an independent survival analysis, VSIG4 and CD163 remained significant after false discovery rate correction, and the signature raised the concordance index over clinical covariates alone from 0.615 to 0.686Enrichment analysis implicated PI3K-Akt, MAPK, focal adhesion, hypoxia, apoptosis, and p53 signaling pathways. This framework couples improved prognostic estimation with biological interpretability supporting multi-omics biomarker discovery in ovarian cancer.

Kosar Fateh, S. Sathipati · 0 citations
Jul 2026

ProphDR: An Interpretable Deep Learning Model for Predicting Cancer Drug Response via Multi-Omics and Cross-Attention Mechanisms.

ProphDR is an interpretable deep learning framework that integrates multiomics data and drug structural information using a hierarchical attention mechanism, and generates biologically interpretable attention maps that highlight key pharmacophores and resistance-related genes consistent with established mechanisms in NSCLC and BRCA.

Yundian Zeng, Qing Ye, Jike Wang et al. · 0 citations
Open access Jul 2026

Essentiality-driven prediction of anticancer drug responses in preclinical and clinical contexts

Summary Precision oncology relies on tumor molecular profiles to predict drug responses. Instead of using conventional molecular features directly, we construct predictive signatures based on gene essentiality. Here, we present DrGee, an essentiality-centered platform that infers drug sensitivity solely from gene expression profiles. The built-in DeepEEAA model integrates gene expression, gene essentiality, drug-protein affinity, and drug-gene associations to quantitatively predict IC50 values. DeepEEAA achieved competitive predictive performance on independent cell line datasets (R2 = 0.764; MSE = 0.9345), outperforming recent benchmark deep learning methods. DrGee prioritized four candidate drugs for the 95-D lung cancer cell line, among which BI-97C1 and trimetrexate were validated by in vitro assays and mouse xenograft experiments. Robust predictive performance was further confirmed in OVCAR8 ovarian cancer cells. In TCGA cohorts, essentiality-driven predictions stratified patients with significantly different overall survival outcomes (AUC-PR = 0.825), highlighting the translational potential of DrGee.

Hongtu Cui, Xiaohui Du, Hai-Xia Guo et al. · 0 citations
Open access Jul 2026

A Comparative Analysis of Machine Learning Models for Cancer Types Classification Using RNA-Seq Gene Expression Data

Cancer is a significant global health concern, and scientists must set the right tumor identifier for accurate diagnosis and personalized treatment plans. While RNA-Seq gene expression data provides critical molecular information, it has two major challenges that can affect ML systems and hinder their performance: its large dimensionality and a propensity for class imbalance. This study offers a comprehensive and data-driven comparison of five machine learning classifiers that demonstrate superior performances in multi-class cancer diagnosis using the RNA-Seq-based TCGA dataset with Random Forest, Support Vector Machine (SVM), and Gradient Boosting k-Nearest Neighbors (k-NN) and Multilayer Perceptron (MLP). Mutual information was used to select important features that reduce the dimensions of the data, and then sensitivity analysis showed the successful performance of the method. To overcome these class imbalance problems, the Synthetic Minority Over-sampling Technique (SMOTE) method was used. The model performance was evaluated using a comprehensive testing framework that integrated 5-fold cross-validation with several evaluation metrics such as balanced accuracy, precision, recall, and F1-score, and confusion matrices and ROC curves. The research group performed ablation studies to better understand the individual effects of correct feature selection and SMOTE on their process. These outcomes show that the MLP classifier reached the highest balanced accuracy (0.9917) after resolving methodological avarice. The assessment demonstrates that SMOTE has a significant impact on improving classification results of minority classes due to its positive effects on recall and F1 score metrics. This study provides insights into selected gene features associated with known oncogenic pathways and, therefore, advanced biological interpretations per se of the obtained results. The work develops a framework that is reproducible and allows multi-metric evaluation of ML-based cancer classification, but exposes two key issues leading to data leakage and validation failures. Our research shows how the promise of AI-driven precision medicine can help build approaches to automated and accurate cancer typing.

Peter Makieu, Sahr Foday, Alfred Santigie Turay · 0 citations
Open access Jul 2026

Multi-omics fusion with machine learning enables robust prediction of treatment response in ovarian cancer for precision population health.

Inter-patient heterogeneity complicates predicting treatment response in ovarian cancer (OC). We developed OMICS-FUSE, an early-fusion multi-omics predictive model integrating proteomic, transcriptomic, and methylomic data from OC patients, evaluated across five machine learning algorithms with SHapley Additive exPlanations (SHAP) and experimental validation. The early-fusion Random Forest model achieved excellent predictive accuracy (AUC = 0.939, accuracy = 0.896, F1 = 0.939), with performance comparable to or surpassing that of the best-performing single-omics models. Nevertheless, the multi-omics framework yielded superior balance across accuracy and F1 score. SHAP analysis identified key determinants of treatment response, including CLEC2A, MYH4, and methylation of SYT12_1, with functional enrichment implicating immune regulation, metabolic pathways, and drug resistance signaling. Experimental validation confirmed six hub genes (CASP8, AQP8, CAV1, FN1, CREB1, KDR), exhibiting expression patterns associated with drug resistance, immune regulation, and prognosis. This multi-omics machine learning model enables robust, interpretable prediction, uncovering molecular signatures for therapeutic stratification and precision oncology in OC.

Jie Chen, Tianshi Mao, Yu Yang et al. · 0 citations