Skip to content
Open access

Integration of regression models and MD simulations for virtual screening of natural compounds: Identification of novel hit scaffolds against SphK1 proteins

Jul 2026 · Journal of King Saud University: Science · Vol 38, pp. 18632025 · 0 citations · 37 references

TL;DR

A strong virtual screening pipeline involving ML with support of molecular docking, MD simulations, and free energies analysis to discover potential virtual Sphk1 inhibitors is presented.

Abstract

Sphingosine kinase 1 (Sphk1) has emerged as a crucial target in the signalling pathways triggered by cancer, making it a current focus in the development of anticancer drugs. This paper used a combined computational method involving machine learning-driven regression models and molecular modelling methods to discover new Sphk1 inhibitors. Several multiple regression models, such as random forest (RF), support vector machine with radial basis function (SVM-RBF), and gradient boosting machine (GBM), k-nearest neighbors (KNN) and elastic-net regularized generalized linear model (GLMNET) were trained using combined descriptors and molecular fingerprint to make a prediction about the new scaffolds and their potency against the Sphk1 protein. The best-performing models were utilised for virtually screening by using the natural product library (TargetMol ∼ 4533 compounds), and the identified candidate molecules were allowed to perform molecular docking to determine their binding interaction with the Sphk1 protein. Upon close inspection, we have found a total six best hits, which have an estimated binding affinity much higher than that of the co-crystallised ligand. The stability and binding behaviour of the six best compounds were further investigated by performing 200 ns molecular dynamics (MD) simulations. Based on docking scores and dynamic stability, two hit candidates (T5711 and T6670) were selected for further analysis. These included MM/PBSA (molecular mechanics/Poisson-Boltzmann surface area) calculations to estimate binding free energies, as well as principal component analysis (PCA) and free energy landscape (FEL) evaluations to assess the conformational flexibility and thermodynamic stability of the protein-ligand complexes. This paper presents a strong virtual screening pipeline involving ML with support of molecular docking, MD simulations, and free energies analysis to discover potential virtual Sphk1 inhibitors. The results provide a platform, on and further additional experimentation and validation/optimisations can be done by medicinal chemists that will lead to the production of effective anticancer agents against the Sphk1 Protein. Graphical abstract showing workflow adopted for identification of new potential natural hits against Sphk1 protein by integrating ML-driven regression and virtual screening-based strategies.

Read PDF

Similar papers

Open access Jul 2026

Machine learning–driven identification of PIM2 kinase inhibitors through QSAR modeling and molecular dynamics simulations

The proto-oncogene serine/threonine kinase PIM2 is a critical regulator of cell proliferation, survival, and tumor progression and represents an attractive therapeutic target for several cancers. In this study, an integrated machine learning–guided computational pipeline was developed to identify potential PIM2 inhibitors by combining quantitative structure–activity relationship (QSAR) modeling, virtual screening, molecular docking, molecular dynamics (MD) simulations, and pharmacokinetic prediction. Bioactivity data for PIM2 inhibitors were retrieved from the ChEMBL database, yielding 5953 compounds. After data cleaning, structural standardization, and removal of duplicates and invalid entries, a curated dataset of 1584 compounds was obtained for QSAR modeling. To address dataset imbalance, the Synthetic Minority Oversampling Technique (SMOTE) was applied before model development. Twelve molecular fingerprint descriptors were generated and used to construct 180 QSAR models using five machine learning algorithms, including Random Forest (RF), Extreme Gradient Boosting (XGBoost), Support Vector Regression (SVR), k-Nearest Neighbors (KNN), and Multilayer Perceptron (MLP). Among these models, the Random Forest–fingerprint model demonstrated the best predictive performance, achieving a mean R2 of 0.971 with low prediction errors (RMSE = 0.271; MAE = 0.125) across training, testing, and cross-validation datasets. The optimized model was subsequently applied to virtual screening of multiple chemical libraries, including FDA-approved drugs, natural product databases, and commercial compound collections. Several promising candidates were identified, including TCMBANKIN000009 (emetine), Amb28533044 (4,6′-Anhydrooxysporidinone), NPC170963 (Lysophosphatidylcholine (15:0)), NPC262768 (Endosulfan), and NPC469603 (8-hydroxyircinialactam A). Molecular docking showed that these compounds bind within the ATP-binding pocket of PIM2 kinase, forming interactions with key residues such as Lys62, Asp125, Asp128, and Glu168. Subsequent molecular dynamics simulations confirmed the stability of selected complexes, demonstrating reduced residue fluctuations, stable protein compactness, and persistent intermolecular interactions during the simulation. Furthermore, ADMET prediction suggested favorable pharmacokinetic and toxicity profiles for several compounds. Collectively, these findings highlight the potential of the identified molecules as promising PIM2 inhibitor candidates, providing valuable leads for future experimental validation and anticancer drug development.

A. Fahira, M. Shahab, Zaheer Ud Din et al. · 0 citations
Open access Jul 2026

Explainable Machine Learning-Guided Virtual Screening and Molecular Docking for Identification of Novel FYN Kinase Inhibitors

FYN kinase is a non-receptor protein tyrosine kinase involved in various cancers and neurodegenerative diseases; however, no selective FYN inhibitor has been approved yet. Here we introduce the explainable Machine Learning (ML) coupled with virtual screening and Molecular Docking (MD) pipeline for fast prediction of new FYN kinase inhibitors. In this study, we constructed the training set of 906 molecules active against FYN kinase from the ChEMBL database. Molecules were encoded with Extended-Connectivity Fingerprints (ECFP4). The classification models Random Forest (RF) and eXtreme Gradient Boosting (XGBoost) were developed, and the latter showed the better performance in test (AUC=0.8118) and 5-fold cross-validation (AUC=0.8297). Based on the SHapley Additive exPlanations (SHAP) values obtained via TreeExplainer, nitrogen-containing heterocycles and hydrogen bond acceptors have been identified as the most important molecular substructures. Using the optimal XGBoost classifier, screening of 2,000 approved drugs has been performed, resulting in 470 hit molecules (23.5% hit rate). Five best molecules were further submitted to the MD procedure using AutoDock Vina to dock to FYN kinase domain (PDB RCSB: 2DQ7), showing binding energies in the interval of -9.57 to -6.32 kcal/mol. Dasatinib Anhydrous (CHEMBL1421) was the second strongest binder (-8.49 kcal/mol), effectively interacting with the ATP binding site. Although CHEMBL1171837 was the strongest binder (-9.57 kcal/mol), it was caught in the ADMET profiling. According to ADMET profiling, the top one inhibitor (CHEMBL1421) satisfies Lipinski’s rule of five and Veber rules. Analysis of hydrogen bond and hydrophobic interactions revealed hydrogen bonding with ASP148, LYS39, and ASN86 and hydrophobic interactions with ALA147, ILE80, and GLY88. Validation by self-docking procedure (self-docking or STS) showed low Root Mean Square Deviation (RMSD)<2.0 Å with a binding affinity of -11.53 kcal/mol. This work highlights how explainable ML can be used in combination with structure-based docking to expedite the drug discovery process against FYN kinase and can be applied to other kinase targets.

Ahmet Turan Demir · 4 citations
Open access Aug 2026

From Descriptor Learning to Binding Stability: An Explainable Machine Learning Pipeline for EGFR Double-Mutant Inhibitor Discovery

An integrated computational workflow combining explainable machine learning, virtual screening, molecular dynamics simulations, and binding free-energy calculations to identify novel inhibitors of this drug-resistant EGFR variant may support the development of new therapeutic strategies for overcoming resistance in EGFR-driven cancers.

Jurica Novak · 0 citations
Jul 2026

Modeling Structure-Activity Relationships with Machine Learning to Identify DPP4 Inhibitors as potential Therapeutics for Type 2 Diabetes

The development of potent dipeptidyl peptidase-4 (DPP4) inhibitors remains a promising therapeutic strategy for the management of type 2 diabetes mellitus (T2DM). In the present study, an integrated computational workflow incorporating machine learning-based quantitative structure-activity relationship (QSAR) modeling, ligand-based virtual screening, molecular docking, molecular dynamics (MD) simulations, and binding free energy calculations was employed to identify novel DPP4 inhibitors. A curated dataset of experimentally validated DPP4 inhibitors was obtained from the ChEMBL database and subjected to systematic preprocessing and molecular descriptor generation. Several machine learning regression algorithms were initially evaluated to identify the most suitable predictive models. The best-performing tree-based algorithms were subsequently optimized and combined using Ridge Stacking and Weighted Average ensemble strategies. Among the developed models, the optimized Ridge Stacking ensemble demonstrated the highest predictive performance, achieving an R 2 of 0.746, an RMSE of 0.819, and a Pearson correlation coefficient of 0.864, indicating strong predictive accuracy and good generalization capability. The robustness of the model was further confirmed through 10-fold cross-validation, bootstrap validation, residual analysis, and applicability domain assessment. The validated ensemble model was then used to screen 95 compounds identified through ligand-based virtual screening. Among these candidates, CP20 exhibited the highest predicted pIC 50 value and was selected for further evaluation together with the reference inhibitor omarigliptin. Molecular docking, structural interaction fingerprinting, molecular dynamics simulations, and MM/GBSA and MM/PBSA binding free energy analyses demonstrated that CP20 formed stable interactions with key catalytic residues of DPP4 and maintained favorable conformational stability throughout the simulation. Collectively, these findings identify CP20 as a promising lead scaffold for the development of novel DPP4 inhibitors and demonstrate the effectiveness of an ensemble machine learning-guided computational framework for accelerating antidiabetic drug discovery. Experimental validation is warranted to confirm its biological activity and therapeutic potential.

Iqra Anwar, T. Chohan, Drakhshaan et al. · 0 citations
Open access Jul 2026

Machine Learning-Driven Discovery of Novel HER2 Inhibitors Through Integrated Virtual Screening and Molecular Dynamics Simulations

Background: HER2 is a key oncogenic gene in breast cancer, involved in tumor progression, metastasis, and therapeutic resistance. This study aimed to find new HER2 inhibitors using a hybrid of machine learning (ML) and structure-based virtual screening (VS), combined with molecular dynamics (MD) simulations on various scaffolds. Methods: Four supervised molecular fingerprint classification models were trained on a dataset of 10,000 validated compounds from ChEMBL. Random Forest was the top model for screening a large compound library. Selected compounds underwent molecular docking in the HER2 ATP binding site, ADMET, drug likeness, toxicity analysis, and 200 ns MD simulations. Methods like PCA, FEL, hydrogen-bond analysis, DCCM, RDF, salt-bridge analysis, and MM/PBSA were used to assess binding stability. Results: Virtual screening identified three compounds, CHMEBL193865 (Lead-1), CHMEBL46740 (Lead-2), and CHMEBL151318 (Lead-3)—with better binding affinity and interaction profiles than the reference inhibitor. MD simulations showed stable protein–ligand complexes with RMSD values of 2.32–2.76 Å. Among these, Lead-2 was the most structurally stable, and Lead-1 had the most favorable binding free energy. All three compounds showed good drug likeness, ADMET properties, and low predicted toxicity. Conclusions: These findings support further in vitro and in vivo testing for developing new therapeutics against HER2-overexpressing breast cancer, highlighting two scaffolds with promising lead optimization potential.

Alhumaidi B. Alabbas, Safar M. Alqahtani · 0 citations
Open access Aug 2026

Discovery of a potent TDP1 inhibitor through machine learning-driven predictive modeling combined with structure-based virtual screening and experimental validation

Tyrosyl-DNA phosphodiesterase I (TDP1) repairs topoisomerase I (TOP1)–mediated DNA damage and is a promising anticancer target, particularly in combination with TOP1 inhibitors. However, the discovery of potent and drug-like TDP1 inhibitors remains challenging due to the limited structural diversity of known active compounds. Here, we developed an integrated computational framework combining machine learning (ML), deep learning (DL), and structure-based docking with experimental validation. A curated dataset of 2040 compounds (857 active, 1183 inactive) was assembled and analyzed by scaffold composition. A total of 40 binary classification models were constructed using six ML algorithms and a deep neural network (DNN), each paired with five molecular fingerprint representations, along with five graph neural network architectures (GCN, GAT, MPNN, AttentiveFP, and FPGNN). The SVM::RDKitDes model performed best (AUC = 0.89, F1 = 0.78, BA = 0.80), with robustness confirmed by Y-scrambling and randomized-split analyses, and SHAP analysis identified 20 key descriptors of TDP1 inhibition. The model was deployed as a web application (http://drugpred.top:5050) and standalone desktop applications (.exe) are available at https://github.com/zenghuang8006/TDP1-inhibitor-prediction. The validated model was applied to screen 201 231 compounds, followed by drug-likeness filtering and hierarchical docking, yielding 16 candidates. Biological evaluation identified compound AO65 as a potent TDP1 inhibitor (IC50 = 0.80 ± 0.02 µM), and quantum chemical calculations and docking elucidated its electronic properties and binding within the catalytic domain. This work demonstrates the value of integrating ML-driven prediction with structure-based approaches and identifies AO65 as a promising lead for further TDP1-focused investigation.

Huang Zeng, Manyi Zhang, Bo Qiu et al. · 0 citations