Aug 2026· International Journal of Molecular Sciences· Vol 27· 0 citations· 72 references
Medicine
TL;DR
An integrated computational workflow combining explainable machine learning, virtual screening, molecular dynamics simulations, and binding free-energy calculations to identify novel inhibitors of this drug-resistant EGFR variant may support the development of new therapeutic strategies for overcoming resistance in EGFR-driven cancers.
Abstract
Drug resistance arising during cancer development and progression remains a major challenge in the treatment of epidermal growth factor receptor (EGFR)-driven tumors, particularly those harboring the clinically relevant T790M/L858R double mutation. In this study, we developed an integrated computational workflow combining explainable machine learning, virtual screening, molecular dynamics simulations, and binding free-energy calculations to identify novel inhibitors of this drug-resistant EGFR variant. An XGBoost regression model was trained using scaffold-aware cross-validation, Bayesian hyperparameter optimization, and sequential feature selection, resulting in a compact model based on 16 molecular descriptors. The model demonstrated robust predictive performance on external validation data, while SHAP analysis identified descriptors related to the local electronic environment, fragment distribution, and molecular topology as the primary contributors to activity prediction. The optimized model was subsequently applied to screen compounds from the Enamine REAL database. Top-ranked candidates were evaluated using explicit-solvent molecular dynamics simulations and MM/GBSA binding free-energy calculations. Several compounds formed stable protein–ligand complexes and maintained key interactions with residues known to be important for EGFR inhibition, including Lys745, Met790, and Leu718. These results demonstrate that the proposed workflow can efficiently prioritize computational candidates of drug-resistant EGFR mutants and may support the development of new therapeutic strategies for overcoming resistance in EGFR-driven cancers.
The proto-oncogene serine/threonine kinase PIM2 is a critical regulator of cell proliferation, survival, and tumor progression and represents an attractive therapeutic target for several cancers. In this study, an integrated machine learning–guided computational pipeline was developed to identify potential PIM2 inhibitors by combining quantitative structure–activity relationship (QSAR) modeling, virtual screening, molecular docking, molecular dynamics (MD) simulations, and pharmacokinetic prediction. Bioactivity data for PIM2 inhibitors were retrieved from the ChEMBL database, yielding 5953 compounds. After data cleaning, structural standardization, and removal of duplicates and invalid entries, a curated dataset of 1584 compounds was obtained for QSAR modeling. To address dataset imbalance, the Synthetic Minority Oversampling Technique (SMOTE) was applied before model development. Twelve molecular fingerprint descriptors were generated and used to construct 180 QSAR models using five machine learning algorithms, including Random Forest (RF), Extreme Gradient Boosting (XGBoost), Support Vector Regression (SVR), k-Nearest Neighbors (KNN), and Multilayer Perceptron (MLP). Among these models, the Random Forest–fingerprint model demonstrated the best predictive performance, achieving a mean R2 of 0.971 with low prediction errors (RMSE = 0.271; MAE = 0.125) across training, testing, and cross-validation datasets. The optimized model was subsequently applied to virtual screening of multiple chemical libraries, including FDA-approved drugs, natural product databases, and commercial compound collections. Several promising candidates were identified, including TCMBANKIN000009 (emetine), Amb28533044 (4,6′-Anhydrooxysporidinone), NPC170963 (Lysophosphatidylcholine (15:0)), NPC262768 (Endosulfan), and NPC469603 (8-hydroxyircinialactam A). Molecular docking showed that these compounds bind within the ATP-binding pocket of PIM2 kinase, forming interactions with key residues such as Lys62, Asp125, Asp128, and Glu168. Subsequent molecular dynamics simulations confirmed the stability of selected complexes, demonstrating reduced residue fluctuations, stable protein compactness, and persistent intermolecular interactions during the simulation. Furthermore, ADMET prediction suggested favorable pharmacokinetic and toxicity profiles for several compounds. Collectively, these findings highlight the potential of the identified molecules as promising PIM2 inhibitor candidates, providing valuable leads for future experimental validation and anticancer drug development.
A. Fahira, M. Shahab, Zaheer Ud Din et al.· Journal of Genetic Engineeri...· 0 citations
Background: HER2 is a key oncogenic gene in breast cancer, involved in tumor progression, metastasis, and therapeutic resistance. This study aimed to find new HER2 inhibitors using a hybrid of machine learning (ML) and structure-based virtual screening (VS), combined with molecular dynamics (MD) simulations on various scaffolds. Methods: Four supervised molecular fingerprint classification models were trained on a dataset of 10,000 validated compounds from ChEMBL. Random Forest was the top model for screening a large compound library. Selected compounds underwent molecular docking in the HER2 ATP binding site, ADMET, drug likeness, toxicity analysis, and 200 ns MD simulations. Methods like PCA, FEL, hydrogen-bond analysis, DCCM, RDF, salt-bridge analysis, and MM/PBSA were used to assess binding stability. Results: Virtual screening identified three compounds, CHMEBL193865 (Lead-1), CHMEBL46740 (Lead-2), and CHMEBL151318 (Lead-3)—with better binding affinity and interaction profiles than the reference inhibitor. MD simulations showed stable protein–ligand complexes with RMSD values of 2.32–2.76 Å. Among these, Lead-2 was the most structurally stable, and Lead-1 had the most favorable binding free energy. All three compounds showed good drug likeness, ADMET properties, and low predicted toxicity. Conclusions: These findings support further in vitro and in vivo testing for developing new therapeutics against HER2-overexpressing breast cancer, highlighting two scaffolds with promising lead optimization potential.
Alhumaidi B. Alabbas, Safar M. Alqahtani· Pharmaceuticals· 0 citations
Mutation-induced drug resistance is a major contributor to the failure of targeted cancer therapies, particularly in tumors driven by mutations in the KRAS oncogene. Although covalent inhibitors effectively target KRAS G12C, secondary mutations such as G12C/Y96C, G12C/Y96S, and G12C/Y96D confer resistance despite leaving the covalent attachment site intact.
To investigate the conformational basis of this resistance, we developed a computational framework integrating molecular dynamics (MD)-derived structural, energetic, thermodynamic, and contact-based descriptors with machine learning. All simulations were performed in the apo (unbound) state; the results therefore reflect conformational and solvent-exposure correlates associated with resistant mutants rather than inhibitor-specific resistance mechanisms. Molecular descriptors extracted from MD simulations of treatment-sensitive and treatment-resistant KRAS systems were used to train logistic regression, random forest, support vector machine, and Bayesian network classifiers. To address the correlated nature of MD-derived conformers, model performance was evaluated using a system-independent mixed held-out validation scheme, and univariate analysis employed clustering-aware statistical approaches.
Cross-referencing machine learning feature importance rankings with within-resistant-group variability testing revealed an important distinction between features reflecting mutant-identity-specific variation and those consistent with a shared resistance phenotype. Residue-level descriptors: solvent-accessible surface area variability at E62 and H95, Lennard-Jones 1,4 interaction energy, and root mean square fluctuation at M72 and H95 were both consistently discriminative across validation schemes and statistically consistent across all three resistant mutants, suggesting that they represent potential apo-state conformational and solvent-exposure signatures associated with resistant KRAS mutants.
Our proof-of-concept workflow may inform the design of inhibitors targeting secondary KRAS resistance mutations, pending validation in additional structurally independent mutant systems.
Katarzyna Mizgalska, Konstancja Urbaniak, Denis Imbody et al.· Frontiers in Chemical Biolog...· 0 citations
The development of potent dipeptidyl peptidase-4 (DPP4) inhibitors remains a promising therapeutic strategy for the management of type 2 diabetes mellitus (T2DM). In the present study, an integrated computational workflow incorporating machine learning-based quantitative structure-activity relationship (QSAR) modeling, ligand-based virtual screening, molecular docking, molecular dynamics (MD) simulations, and binding free energy calculations was employed to identify novel DPP4 inhibitors. A curated dataset of experimentally validated DPP4 inhibitors was obtained from the ChEMBL database and subjected to systematic preprocessing and molecular descriptor generation. Several machine learning regression algorithms were initially evaluated to identify the most suitable predictive models. The best-performing tree-based algorithms were subsequently optimized and combined using Ridge Stacking and Weighted Average ensemble strategies. Among the developed models, the optimized Ridge Stacking ensemble demonstrated the highest predictive performance, achieving an R
2
of 0.746, an RMSE of 0.819, and a Pearson correlation coefficient of 0.864, indicating strong predictive accuracy and good generalization capability. The robustness of the model was further confirmed through 10-fold cross-validation, bootstrap validation, residual analysis, and applicability domain assessment. The validated ensemble model was then used to screen 95 compounds identified through ligand-based virtual screening. Among these candidates, CP20 exhibited the highest predicted pIC
50
value and was selected for further evaluation together with the reference inhibitor omarigliptin. Molecular docking, structural interaction fingerprinting, molecular dynamics simulations, and MM/GBSA and MM/PBSA binding free energy analyses demonstrated that CP20 formed stable interactions with key catalytic residues of DPP4 and maintained favorable conformational stability throughout the simulation. Collectively, these findings identify CP20 as a promising lead scaffold for the development of novel DPP4 inhibitors and demonstrate the effectiveness of an ensemble machine learning-guided computational framework for accelerating antidiabetic drug discovery. Experimental validation is warranted to confirm its biological activity and therapeutic potential.
Iqra Anwar, T. Chohan, Drakhshaan et al.· Journal of Computational Bio...· 0 citations
FYN kinase is a non-receptor protein tyrosine kinase involved in various cancers and neurodegenerative diseases; however, no selective FYN inhibitor has been approved yet. Here we introduce the explainable Machine Learning (ML) coupled with virtual screening and Molecular Docking (MD) pipeline for fast prediction of new FYN kinase inhibitors. In this study, we constructed the training set of 906 molecules active against FYN kinase from the ChEMBL database. Molecules were encoded with Extended-Connectivity Fingerprints (ECFP4). The classification models Random Forest (RF) and eXtreme Gradient Boosting (XGBoost) were developed, and the latter showed the better performance in test (AUC=0.8118) and 5-fold cross-validation (AUC=0.8297). Based on the SHapley Additive exPlanations (SHAP) values obtained via TreeExplainer, nitrogen-containing heterocycles and hydrogen bond acceptors have been identified as the most important molecular substructures. Using the optimal XGBoost classifier, screening of 2,000 approved drugs has been performed, resulting in 470 hit molecules (23.5% hit rate). Five best molecules were further submitted to the MD procedure using AutoDock Vina to dock to FYN kinase domain (PDB RCSB: 2DQ7), showing binding energies in the interval of -9.57 to -6.32 kcal/mol. Dasatinib Anhydrous (CHEMBL1421) was the second strongest binder (-8.49 kcal/mol), effectively interacting with the ATP binding site. Although CHEMBL1171837 was the strongest binder (-9.57 kcal/mol), it was caught in the ADMET profiling. According to ADMET profiling, the top one inhibitor (CHEMBL1421) satisfies Lipinski’s rule of five and Veber rules. Analysis of hydrogen bond and hydrophobic interactions revealed hydrogen bonding with ASP148, LYS39, and ASN86 and hydrophobic interactions with ALA147, ILE80, and GLY88. Validation by self-docking procedure (self-docking or STS) showed low Root Mean Square Deviation (RMSD)<2.0 Å with a binding affinity of -11.53 kcal/mol. This work highlights how explainable ML can be used in combination with structure-based docking to expedite the drug discovery process against FYN kinase and can be applied to other kinase targets.
Ahmet Turan Demir· Intelligent Systems Research...· 4 citations
Glioblastoma (GB) is an aggressive and lethal brain tumor characterized by high mortality and poor prognosis. Amplification and mutation of the epidermal growth factor receptor (EGFR) gene are key drivers of GB progression, highlighting EGFR as a promising therapeutic target. This study aimed to identify novel small-molecule EGFR inhibitors through a pharmacophore-guided computational screening approach. A pharmacophore model was generated using the co-crystal ligand (PDB: HYZ) as a template. This model was employed to screen the Specs database comprising approximately 280,000 compounds. Ligand-based virtual screening yielded 323 hits that matched the pharmacophore query. These compounds were subjected to molecular docking against the EGFR active site using the Glide standard precision protocol, applying a binding affinity threshold of -9 kcal/mol. Eleven compounds with favorable docking scores were further evaluated through ADMET (adsorption, distribution, metabolism, excretion, toxicity) profiling, and eight top candidates were subjected to molecular dynamics simulations to assess binding stability. The pharmacophore-based screening and docking analysis identified eleven promising EGFR-binding compounds, of which eight demonstrated optimal ADMET characteristics and stable interactions within the active site during molecular dynamics simulations. The selected molecules exhibited strong binding affinities and favorable conformational stability, suggesting their potential efficacy as EGFR inhibitors. This study identified and characterized novel small molecules with high potential for EGFR inhibition in glioblastoma. These findings provide a foundation for further experimental validation and may contribute to the development of targeted therapies for GB.
M. Moulay, M. Mahmoud, Reem M. Farsi et al.· Journal of King Saud Univers...· 0 citations