Skip to content
Open access

An Explainable Machine Learning-Based QSAR Framework for Predicting Thrombin Inhibitory Activity

Sep 2026 · Pharmaceuticals · Vol 19, pp. 1411 · 0 citations · 52 references
Medicine

TL;DR

An explainable, assay-aware, and leakage-safe machine learning framework for predicting thrombin-inhibitory activity provides useful predictions within the represented chemical space and measurable generalization for unseen scaffolds.

Abstract

Background/Objectives: Public thrombin bioactivity records contain heterogeneous endpoints, replicate measurements, related chemical series, and potentially reactive compounds that may bias quantitative structure–activity relationship models. In this study, we developed an explainable, assay-aware, and leakage-safe machine learning framework for predicting thrombin-inhibitory activity. Methods: Exact Ki and IC50 records for human thrombin (CHEMBL204) were standardized, converted to pActivity, and aggregated using predefined criteria. The final dataset comprised 5189 unique compounds represented by development-filtered Mordred descriptors and Morgan fingerprint. The models were optimized using development-only out-of-fold validation and evaluated using random and scaffold-disjoint held-out tests. Results: In the random-split analysis, ConsensusAll achieved an out-of-fold R2 of 0.7588 and a held-out test R2 of 0.7480, with an RMSE of 0.7360, MAE of 0.5307, and concordance correlation coefficient of 0.8540. In the scaffold-disjoint locked test, ConsensusTop5 achieved R2 = 0.5831, RMSE = 0.9484, and MAE = 0.7290, respectively. The applicability domain covered 92.68% of the locked test compounds and yielded R2 = 0.6041. One hundred Y-randomization runs produced a mean R2 of −0.1233 (empirical p = 0.0099). Conclusions: The framework provides useful predictions within the represented chemical space and measurable generalization for unseen scaffolds. This supports compound prioritization, although prospective biochemical validation remains necessary.

Read PDF

Similar papers

Open access Sep 2026

Development of an Interpretable QSAR Model for Predicting Coagulation Factor XIIa Inhibitors Using Ensemble Machine Learning

Background/Objective: Activated coagulation factor XII (FXIIa) is a component of the contact activation pathway and a pharmacologically relevant target in contact-system-associated processes. In this study, scaffold-aware and interpretable machine-learning QSAR models were developed for human FXIIa activity. Methods: B...

Ali Onur Kaya, Mert Can Emre · 0 citations
Sep 2026

Interpretable Machine Learning for Structure–Activity Relationship Modeling of Cyclooxygenase-1 Inhibitors

An interpretable computational framework is developed for modeling COX-1 inhibitor activity and characterizing prioritized compounds at the molecular level by integrating QSAR modeling, explainable artificial intelligence, molecular docking, molecular dynamics, and MM/GBSA calculations to investigate COX-1 inhibitors.

Hamid Bouseber, Mustapha Cherkaoui · 0 citations
Sep 2026

Explainable Machine Learning for Predicting Antibacterial Activity of Natural Products

Natural products are a valuable source of bioactive compounds with considerable potential for antibacterial drug discovery. However, experimental screening of large natural-product collections is often time-consuming and resource-intensive. In this study, an explainable machine-learning framework was developed to predi...

Deepthi Goteti · 0 citations
Open access 2026

EXPLAINABLE MACHINE LEARNING FOR PREDICTING MOLECULAR TOXICITY FROM PHYSICOCHEMICAL AND STRUCTURAL PROPERTIES

The prediction of molecular toxicity is becoming more and more critical for effective chemical safety assessment, drug development and environmental risk assessment. In this study, an explainable machine learning model for predicting molecular toxicity was developed using physicochemical descriptors and structural prop...

S. Kulkarni · 0 citations
Open access Oct 2026

Pan-Screening the Structural Predictability of Drug Toxicity: Validation of Molecular Fingerprints

Objective To systematically evaluate the predictability of drug chemical structures against the complete set of MedDRA Preferred Terms (PTs), construct a comprehensive structure– adverse reaction landscape, validate the utility of molecular fingerprint representations in pan-screening tasks, and establish quantitative...

Wei Jiang, Xiang'e Jian, Jie-Kai Hu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.