Aug 2026· PLoS ONE· Vol 21, pp. e0345854· 0 citations· 65 references
Medicine
TL;DR
This study systematically benchmark six ranking loss functions, including state-of-the-art listwise methods, and five types of molecular representations across two large-scale drug screening datasets, CTRP and PRISM, to demonstrate that listwise loss functions such as LambdaLoss and LambdaRank consistently excel in both early and overall ranking quality.
Abstract
Learning to Rank (LeToR) methods have gained increasing attention in drug response prediction, offering a direct way to prioritize effective treatments for cancer cell lines. In this study, we systematically benchmark six ranking loss functions, including state-of-the-art listwise methods, and five types of molecular representations across two large-scale drug screening datasets, CTRP and PRISM. Using high-dimensional gene expression profiles and various drug fingerprints and descriptors, we evaluated models under multiple validation setups and ranking metrics. Our results demonstrate that listwise loss functions such as LambdaLoss and LambdaRank consistently excel in both early and overall ranking quality. Additionally, combining molecular fingerprints with physicochemical descriptors yielded improved performance. A novel attention-based mechanism and a modified version of RankingSHAP were integrated to enhance interpretability, uncovering key genes and substructures aligned with known biological insights. The explainability pipeline successfully distinguished estrogen receptor-positive (ER⁺) and estrogen receptor-negative (ER−) breast cancer subtypes. The model successfully identified critical substructures in docetaxel, an FDA-approved therapy, and triptolide, which is currently undergoing clinical evaluation for breast cancer. These findings are consistent with established structure-activity relationship (SAR) data. Overall, this study presents a comprehensive evaluation framework and underscores the importance of carefully selecting loss functions and feature representations when developing robust and interpretable drug-ranking systems.
Summary Precision oncology relies on tumor molecular profiles to predict drug responses. Instead of using conventional molecular features directly, we construct predictive signatures based on gene essentiality. Here, we present DrGee, an essentiality-centered platform that infers drug sensitivity solely from gene expression profiles. The built-in DeepEEAA model integrates gene expression, gene essentiality, drug-protein affinity, and drug-gene associations to quantitatively predict IC50 values. DeepEEAA achieved competitive predictive performance on independent cell line datasets (R2 = 0.764; MSE = 0.9345), outperforming recent benchmark deep learning methods. DrGee prioritized four candidate drugs for the 95-D lung cancer cell line, among which BI-97C1 and trimetrexate were validated by in vitro assays and mouse xenograft experiments. Robust predictive performance was further confirmed in OVCAR8 ovarian cancer cells. In TCGA cohorts, essentiality-driven predictions stratified patients with significantly different overall survival outcomes (AUC-PR = 0.825), highlighting the translational potential of DrGee.
Hongtu Cui, Xiaohui Du, Hai-Xia Guo et al.· iScience· 0 citations
An integrated computational workflow combining explainable machine learning, virtual screening, molecular dynamics simulations, and binding free-energy calculations to identify novel inhibitors of this drug-resistant EGFR variant may support the development of new therapeutic strategies for overcoming resistance in EGFR-driven cancers.
Jurica Novak· International Journal of Mol...· 0 citations
Accurate prediction of drug sensitivity in cancer cell lines is vital for precision oncology and patient-specific therapies. However, many computational approaches fail to integrate multi-modal biological and chemical features and often struggle with high-dimensional, imbalanced pharmacogenomic data, limiting predictive accuracy and interpretability. To address these challenges, we developed a machine learning framework that integrates pharmacogenomic profiles-including mutation status, copy number alterations, and microsatellite instabil-ity-with molecular fingerprints and descriptors of 85 anticancer drugs, generated using PaDEL from SMILES strings. Data from 40 breast cancer cell lines in the Genomics of Drug Sensitivity in Cancer (GDSC) dataset were employed. A threestage feature selection strategy combining Boruta, mRMR, and XGBoost was applied to reduce drug feature dimensionality while retaining 130 cell line features. Multiple models were trained, and LightGBM, optimized with grid search, class weighting, and 3-fold cross-validation, demonstrated superior performance in handling severe class imbalance (233 sensitive vs. 3167 resistant samples). LightGBM achieved training AUROC $=0.9455$, AUPRC $\boldsymbol{=} \mathbf{0. 5 1 4 8}$, Accuracy $\boldsymbol{=} \mathbf{0. 8 4 1 5}$, F1-score = 0.4481, Recall = 0.9409, and MCC = 0.4732, underscoring its suitability for sparse biomedical datasets. Model interpretation with SHapley Additive exPlanations (SHAP) highlighted BRCA-related features, identifying cnaBRCA25 (not mutated) as a resistance marker and cnaBRCA47 (mutated) as a context-dependent biomarker, consistent with their roles in DNA repair pathways. Overall, this framework demonstrates the value of multi-modal integration and interpretable machine learning in pharmacogenomics. While results are promising, validation on larger and independent cohorts is essential to establish clinical relevance.
D. Kumari, Aiman, Sakshi Singh et al.· Annual International Compute...· 0 citations
Quantitative prediction of inhibitor potency can accelerate early-stage drug discovery. Recently, data-driven approaches have gained widespread interest in drug discovery, as evidenced by a growing number of benchmarking challenges and open competitions. In this context, we developed a machine learning-based methodology that can find the most effective way of predicting IC50 values against ASK1 from SMILES, for "Jump AI(.py) 2025: 3rd AI Drug Discovery Competition", hosted by the Korea Pharmaceutical and Bio-Pharma Manufacturers Association (KPBMA) on the Dacon platform. Applying our methodology achieved the highest overall predictive performance among all participating teams. Beyond this competition setting, we present a compact SMILES-based modeling workflow comprising (i) a pre-trained encoder, (ii) regression models, (iii) data augmentation, and (iv) hyperparameter tuning. We systematically compared molecular representations from sequence- and graph-based models, including ChemBERTa-2 and MolCLR. Across encoder-regressor combinations, ChemBERTa-77 M-MLM embeddings paired with support vector regression (SVR) yielded the strongest predictive performance. Embedding-level mix-up augmentation and SVR hyperparameter tuning further improved predictive performance. Our findings highlight that careful SMILES preprocessing and encoder selection have a critical influence on IC50 values and provide a reproducible benchmark for single-target bioactivity prediction, thus contributing to a more efficient drug discovery process. Scientific Contribution In this study, we propose a machine learning methodology for predicting the IC50 values of ASK1 inhibitors from SMILES representations, with a systematic comparison of molecular encoders and regression models. Our results show that the use of suitable encoder-regressor pairs together with embedding-level mix-up augmentation improves model generalizability without requiring SMILES-level augmentation. This strategy would be particularly useful for settings with imbalanced labels or limited data, and could be applied more broadly to IC50 prediction for other kinase inhibitors.
Ju Hyung Lee, S. Choi, Utku Ozbulak et al.· Journal of Cheminformatics· 0 citations
Identifying synergistic anti-cancer drug combinations is crucial for improving efficacy and reducing toxicity, but exhaustive experimental screening is prohibitively costly. We present SYNAPX, an explainable deep learning framework for drug synergy prediction that integrates chemical features of drug pairs with gene expression profiles of cancer cell lines. Our model combines ECFP6 fingerprints, physicochemical descriptors, toxicophore features, and transcriptomic features into a unified representation, and uses a fully connected neural network to predict continuous synergy scores. We further apply SHAP (SHapley Additive exPlanations) for biological interpretation to quantify feature contributions and explain individual predictions. Evaluated on the drug combination dataset of 23,052 oncology drug-combination samples, the method achieves strong predictive performance, including a ROC AUC of 0.9169, PR AUC of 0.7139, and balanced accuracy of 0.8289 after threshold optimization. Our explanation analysis shows that the most influential features are primarily molecular fingerprints and physicochemical descriptors, and a case study on the Methotrexate-BEZ-235 combination yields explanations consistent with known biological mechanisms.
Yazhini Kalaignan, Yepeng Ding· International Conference on...· 0 citations
ProphDR is an interpretable deep learning framework that integrates multiomics data and drug structural information using a hierarchical attention mechanism, and generates biologically interpretable attention maps that highlight key pharmacophores and resistance-related genes consistent with established mechanisms in NSCLC and BRCA.
Yundian Zeng, Qing Ye, Jike Wang et al.· Journal of Chemical Informat...· 0 citations