Skip to content
Open access

Beyond the Score: Fixed-Budget Benchmarking of Virtual Screening Integration Strategies for Decision-Centric Drug Discovery

Aug 2026 · International Journal of Molecular Sciences · Vol 27 · 0 citations · 33 references
Medicine

TL;DR

Target-level results showed substantial variability in MCS-containing workflows and limited benefits from adding docking without target-specific optimization, and a validated ligand-based predictor or a simple two-method rank-fusion scheme provided the highest observed mean hit recovery without requiring elaborate integration.

Abstract

Virtual screening (VS) workflows often combine structure- and ligand-based methods; however, their value depends on the number of compounds that can be tested. We benchmarked 20 fixed-budget strategies derived from molecular docking (GNINA CNN score), maximum common substructure (MCS) similarity, and a calibrated machine-learning (ML)-QSAR classifier across five pharmacologically diverse targets. Individual methods, best-rank and worst-rank fusion, mean-rank consensus, and sequential funnels were evaluated at 1%, 5%, and 10% library fractions, with every strategy selecting the same number of compounds. ML-QSAR was the strongest standalone method, recovering 47.6%, 81.6%, and 84.4% of actives at the three cutoffs. At the 1% budget, ML-QSAR achieved the highest mean hit recovery (47.6% recall; 99.2% precision). At 5% and 10%, best-rank fusion of QSAR and MCS produced the highest mean recall (83.2% and 86.4%). Among the sequential workflows, QSAR → MCS achieved the highest 1% hit recovery (45.2 ± 3.3% recall), whereas docking-first funnels consistently underperformed under the default, non-optimized conditions evaluated in this study. Target-level results showed substantial variability in MCS-containing workflows and limited benefits from adding docking without target-specific optimization. Under matched assay budgets, a validated ligand-based predictor or a simple two-method rank-fusion scheme provided the highest observed mean hit recovery without requiring elaborate integration.

Read PDF

Similar papers

Open access Sep 2026

Data driven selection of consensus docking pipelines for structure based hit identification

Structure-based virtual screening (SBVS) is a cornerstone of computer-aided drug design, yet its success depends on selecting a combination of docking tools, scoring function (SF), and ranking strategies. MolDockLab addresses this challenge with an automated, data-driven framework that optimizes SBVS workflows for a pr...

Hamza Agha, Y. Ibrahim, Michael Backenköhler et al. · 0 citations
Jul 2026

Real-World Assessment of Machine-Learned Docking Using Bioassay-Derived Benchmarks

This work systematically evaluates the performance of a popular ML-based docking method, DiffDock-Pocket, on high-throughput screening (HTS) data sets derived from the PubChem BioAssay database, a premier source of bioactivity data.

Furyal Ahmed, M. Soellner, Charles L. Brooks · 0 citations
Open access Sep 2026

Docking-score landscapes shape active-learning performance across Vina, Glide, and SILCS

The rapid expansion of large chemical libraries has created a need for virtual screening workflows that are both efficient and accurate. Active learning (AL) offers a scalable strategy by iteratively training surrogate models to prioritize promising compounds and reduce the number of required docking calculations. Howe...

Joseph Chung, Aashish Bhatt, Jacob Ede Levine et al. · 0 citations

Systematic Evaluation of Graph Neural Networks for Ligand-Based Virtual Screening on ChEMBL Datasets

The performance of target-specific, ligand-based virtual screening models is strongly influenced by dataset characteristics, including data availability, class imbalance, and evaluation strategies. In this work, we perform a systematic evaluation of graph neural networks (GNNs) using a ChEMBL-derived dataset spanning...

Haihan Liu, Jiaqi Lin, Ying Fan et al. · 0 citations
Open access Sep 2026

AI-enhanced adaptive virtual screening of large libraries for ligand discovery.

Ultralarge virtual screenings (ULVSs) evaluate billions of molecules for drug discovery but face cost, flexibility and scalability limits. We introduce AdaptiveFlow, an open-source platform that makes ULVSs more accessible, scalable and efficient and supports artificial intelligence (AI) and machine learning (ML) metho...

Domiziana Cecchini, AkshatKumar Nigam, Ming Tang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.