Aug 2026· Diseases of the esophagus· Vol 39· 0 citations
TL;DR
The natural product compounds CNP0456830 and CNP0467494 exhibited the lowest binding free energies for both EGFR and PIK3CA, identifying them as the most promising dual-target inhibitors.
Abstract
Esophageal Cancer: Molecular Biology/Pathology
With the advancement of personalized medicine, multi-target drug development has garnered significant attention, particularly for complex diseases such as cancer. This study aims to identify potential dual-target inhibitors against Epidermal Growth Factor Receptor (EGFR) and Phosphatidylinositol-4,5-bisphosphate 3-kinase catalytic subunit alpha (PIK3CA), two proteins whose aberrant activation is closely associated with tumorigenesis and progression in various cancers.
We collected IC50 values of active compounds for EGFR and PIK3CA from the BindingDB database, which were then standardized to pIC50 values using RDKit. A total of 2048 Extended-Connectivity Fingerprints (ECFPs) were calculated to serve as molecular descriptors. Various machine learning models, including Support Vector Machine (SVM), Decision Tree, Random Forest, Gradient Boosting, K-Nearest Neighbors, and LightGBM, were developed. The optimal model parameters were determined using ten-fold cross-validation and grid search, and model performance was assessed by Mean Absolute Error (MAE), Mean Squared Error (MSE), and the R-squared (R2) value.
The SVM model demonstrated the best performance and was selected to predict activities for both EGFR and PIK3CA.
The natural product compounds CNP0456830 and CNP0467494 exhibited the lowest binding free energies for both EGFR and PIK3CA, identifying them as the most promising dual-target inhibitors. This study offers a new direction and a potential therapeutic strategy for personalized drug design in cancer treatment.
Aurora kinase A (AURKA) is a pivotal driver of malignant progression and poor prognosis in triple-negative breast cancer (TNBC). In this study, we developed a cascaded AI-driven virtual screening pipeline, integrating sequence-based affinity prediction (PSICHIC), equivariant deep learning docking (KarmaDock), and geometric rescoring (DeepDock) to identify novel AURKA inhibitor candidates. From an in-house 160,000-compound screening library assembled from commercially available collections, three leads (compounds 3, 5, and 8) were selected and subsequently validated via HTRF biochemical assays, exhibiting potent enzymatic inhibition with IC50 values of 157 nM, 21.64 nM, and 46.03 nM, respectively. Cell-based assays demonstrated that compound 3 produced stronger short-term cell-growth inhibition in MDA-MB-231 (TNBC) cells compared to clinical benchmarks MLN8237 and CCT241736, whereas compounds 3 and 5 showed cell-growth inhibition in NIH/3T3 cells within the same concentration range as the reference inhibitors. Triplicate 500 ns molecular dynamics simulations supported stable binding modes of the identified leads in the AURKA binding pocket. Additional computational analyses further provided supportive information for subsequent lead optimization. This study provides a transparent and open-source workflow for AI-assisted identification of AURKA-active chemotypes.
Background: Anaplastic Lymphoma Kinase (ALK) is an oncogenic receptor tyrosine kinase implicated in several cancers. Despite the clinical success of ALK inhibitors, acquired resistance continues to drive the search for novel chemotypes. We developed a multiclass machine learning framework to classify ALK inhibitory activity using a curated ChEMBL dataset. Methods: Models were built using 2D molecular descriptors together with MACCS and ECFP4 fingerprints. Three widely used algorithms, Support Vector Machine (SVM), Random Forest (RF), and XGBoost, were applied for model development. Results: RF and XGBoost models demonstrated the best performance, achieving accuracies of ~0.75–0.79 with consistently high ROC–AUC values, particularly for fingerprint-based features. Bemis–Murcko scaffold analysis identified enriched chemotypes and underexplored scaffolds for further prioritization. The validated models were subsequently used to screen the Maybridge library, and compounds predicted to possess potential ALK inhibitory activity were prioritized for further computational evaluation. Applicability-domain filtering confirmed that the selected compounds occupied the predicted ALK inhibitor chemical space across multiple activity classes. The shortlisted compounds were subsequently evaluated by molecular docking to characterize their binding modes and interactions. Three candidate hits (SCR00078, SCR00073, and AW01085) were selected for further evaluation using 500 ns molecular dynamics simulations alongside the reference inhibitor Brigatinib. Simulation analyses revealed stable protein–ligand complexes and reduced conformational fluctuations relative to apo ALK, while MM/PBSA calculations identified SCR00078 and AW01085 as the most favorable binders. Conclusions: This integrated ML-to-simulation workflow prioritizes structurally novel candidate hits with predicted ALK inhibitory activity and provides an effective strategy for scaffold discovery and hit prioritization.
José Zarzuelo Romero, M. López-Viota, M. Haque et al.· Pharmaceuticals· 0 citations
The proto-oncogene serine/threonine kinase PIM2 is a critical regulator of cell proliferation, survival, and tumor progression and represents an attractive therapeutic target for several cancers. In this study, an integrated machine learning–guided computational pipeline was developed to identify potential PIM2 inhibitors by combining quantitative structure–activity relationship (QSAR) modeling, virtual screening, molecular docking, molecular dynamics (MD) simulations, and pharmacokinetic prediction. Bioactivity data for PIM2 inhibitors were retrieved from the ChEMBL database, yielding 5953 compounds. After data cleaning, structural standardization, and removal of duplicates and invalid entries, a curated dataset of 1584 compounds was obtained for QSAR modeling. To address dataset imbalance, the Synthetic Minority Oversampling Technique (SMOTE) was applied before model development. Twelve molecular fingerprint descriptors were generated and used to construct 180 QSAR models using five machine learning algorithms, including Random Forest (RF), Extreme Gradient Boosting (XGBoost), Support Vector Regression (SVR), k-Nearest Neighbors (KNN), and Multilayer Perceptron (MLP). Among these models, the Random Forest–fingerprint model demonstrated the best predictive performance, achieving a mean R2 of 0.971 with low prediction errors (RMSE = 0.271; MAE = 0.125) across training, testing, and cross-validation datasets. The optimized model was subsequently applied to virtual screening of multiple chemical libraries, including FDA-approved drugs, natural product databases, and commercial compound collections. Several promising candidates were identified, including TCMBANKIN000009 (emetine), Amb28533044 (4,6′-Anhydrooxysporidinone), NPC170963 (Lysophosphatidylcholine (15:0)), NPC262768 (Endosulfan), and NPC469603 (8-hydroxyircinialactam A). Molecular docking showed that these compounds bind within the ATP-binding pocket of PIM2 kinase, forming interactions with key residues such as Lys62, Asp125, Asp128, and Glu168. Subsequent molecular dynamics simulations confirmed the stability of selected complexes, demonstrating reduced residue fluctuations, stable protein compactness, and persistent intermolecular interactions during the simulation. Furthermore, ADMET prediction suggested favorable pharmacokinetic and toxicity profiles for several compounds. Collectively, these findings highlight the potential of the identified molecules as promising PIM2 inhibitor candidates, providing valuable leads for future experimental validation and anticancer drug development.
A. Fahira, M. Shahab, Zaheer Ud Din et al.· Journal of Genetic Engineeri...· 0 citations
Interleukin-1 receptor-associated kinase 4 (IRAK4) is one of the IRAK family proteins and plays an important role in the regulation of innate and inflammatory responses. In particular, IRAK4 acts as a key regulator of the Toll-like receptor (TLR) and interleukin-1 receptor (IL-1R) signaling pathways and has attracted attention as a therapeutic target for immune and inflammatory diseases. In this study, an integrated computational approach combining machine learning, molecular docking, and molecular dynamics simulations was applied to identify putative IRAK4 inhibitor candidates. Bioactivity data of IRAK4 were obtained from the ChEMBL and PubChem databases and evaluated for multiple binary classification models. The optimized XGBoost model based on ECFP4 and PubChem fingerprints achieved an ROC-AUC of 0.996 and an average precision (AP) of 0.991 on the independent test set. After that, 20 candidate compounds with high predictive probability score were finally selected through subsequent screening of the DrugBank database. Among them, DB12168 (MK-0557), DB15040 (TP-271), and DB18152 (Zilurgisertib) exhibited favorable binding free energies and stable complex formation with IRAK4 through molecular dynamics simulations and MM-PBSA calculations. Overall, these results demonstrate that approaches incorporating machine learning and structure-based computational analysis can be useful for discovering and prioritizing potential IRAK4 inhibitor candidates.
H. Na, Juwon Park, Jiwon Choi· Current Issues in Molecular...· 0 citations
An integrated computational workflow combining explainable machine learning, virtual screening, molecular dynamics simulations, and binding free-energy calculations to identify novel inhibitors of this drug-resistant EGFR variant may support the development of new therapeutic strategies for overcoming resistance in EGFR-driven cancers.
Jurica Novak· International Journal of Mol...· 0 citations
FYN kinase is a non-receptor protein tyrosine kinase involved in various cancers and neurodegenerative diseases; however, no selective FYN inhibitor has been approved yet. Here we introduce the explainable Machine Learning (ML) coupled with virtual screening and Molecular Docking (MD) pipeline for fast prediction of new FYN kinase inhibitors. In this study, we constructed the training set of 906 molecules active against FYN kinase from the ChEMBL database. Molecules were encoded with Extended-Connectivity Fingerprints (ECFP4). The classification models Random Forest (RF) and eXtreme Gradient Boosting (XGBoost) were developed, and the latter showed the better performance in test (AUC=0.8118) and 5-fold cross-validation (AUC=0.8297). Based on the SHapley Additive exPlanations (SHAP) values obtained via TreeExplainer, nitrogen-containing heterocycles and hydrogen bond acceptors have been identified as the most important molecular substructures. Using the optimal XGBoost classifier, screening of 2,000 approved drugs has been performed, resulting in 470 hit molecules (23.5% hit rate). Five best molecules were further submitted to the MD procedure using AutoDock Vina to dock to FYN kinase domain (PDB RCSB: 2DQ7), showing binding energies in the interval of -9.57 to -6.32 kcal/mol. Dasatinib Anhydrous (CHEMBL1421) was the second strongest binder (-8.49 kcal/mol), effectively interacting with the ATP binding site. Although CHEMBL1171837 was the strongest binder (-9.57 kcal/mol), it was caught in the ADMET profiling. According to ADMET profiling, the top one inhibitor (CHEMBL1421) satisfies Lipinski’s rule of five and Veber rules. Analysis of hydrogen bond and hydrophobic interactions revealed hydrogen bonding with ASP148, LYS39, and ASN86 and hydrophobic interactions with ALA147, ILE80, and GLY88. Validation by self-docking procedure (self-docking or STS) showed low Root Mean Square Deviation (RMSD)<2.0 Å with a binding affinity of -11.53 kcal/mol. This work highlights how explainable ML can be used in combination with structure-based docking to expedite the drug discovery process against FYN kinase and can be applied to other kinase targets.
Ahmet Turan Demir· Intelligent Systems Research...· 4 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduAug 27, 2026
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.