Skip to content
Open access

Chemical Distance-Based Acceleration of Large Library Docking with ChemSTEP

Jul 2026 · ACS Central Science · Vol 12, pp. 1178 - 1194 · 0 citations · 59 references
Medicine

TL;DR

A similarity-based prioritization approach, ChemSTEP, is introduced, and the effective size of a library treated by any prioritization algorithm, Neff is defined, which suggests trillion-molecule libraries might be in reach with this approach.

Abstract

While make-on-demand libraries now span trillions of molecules, full library docking struggles beyond a few billion, motivating prioritization that recovers top-scoring compounds while evaluating only a fraction of a library. Here we introduce a similarity-based prioritization approach, ChemSTEP, and define the effective size of a library treated by any prioritization algorithm, Neff. ChemSTEP docks a representative seed set, selects diverse high-scoring “beacons”, and iteratively traverses the library through cycles of beacon selection, similarity search, and docking. Retrospectively on eight targets, ChemSTEP recovered over 75% of high-scoring compounds while docking less than 5% of a library. We then tested ChemSTEP prospectively against AmpC β-lactamase using a 13.2 billion molecule library. Because AmpC recognizes negatively charged inhibitors, we explicitly docked all 360 million library anions, synthesizing and testing 241 high-ranking ones in parallel to the ChemSTEP 13.2B run. Compared with previous docking of 99 million and 1.7 billion molecules against AmpC, the 13.2 billion library had higher hit-rates (2% vs 25% vs 37%, respectively) and found more potent compounds. Meanwhile, ChemSTEP retrieved 80% of the 241 high-ranking compounds within the first 0.5% docked. Trillion-molecule libraries might be in reach with this approach.

Read PDF

Similar papers

Open access Sep 2026

When DL-Based Prescreening Meets Synthon-Based Docking: Target-Adapting PharmacoNet via MEL-Steered Correction

As chemical libraries expand into the trillions of molecules, Virtual SYNthon Hierarchical Enumeration Screening (V-SYNTHES) has emerged as a leading strategy for making gigascale virtual screening computationally tractable. In V-SYNTHES, a Minimal Enumeration Library (MEL) of chemical fragments is docked against a tar...

Wen-Jin Liu, Yong-Chan Hong, Thomas Ku et al. · 0 citations
Open access Sep 2026

An extension of jamdock-suite for virtual screening from raw chemical datasets.

Virtual screening of chemical libraries has long been a cornerstone for identifying candidate bioactive compounds. Recently, many tools have been developed to make virtual screening more accessible. Among them, the jamdock-suite provides an automated and user-friendly workflow that encompasses the entire process from l...

Shahariar Emon, M. Haque, Iftekhar Alam et al. · 0 citations
Open access Sep 2026

Docking-score landscapes shape active-learning performance across Vina, Glide, and SILCS

The rapid expansion of large chemical libraries has created a need for virtual screening workflows that are both efficient and accurate. Active learning (AL) offers a scalable strategy by iteratively training surrogate models to prioritize promising compounds and reduce the number of required docking calculations. Howe...

Joseph Chung, Aashish Bhatt, Jacob Ede Levine et al. · 0 citations
Open access Sep 2026

Data driven selection of consensus docking pipelines for structure based hit identification

Structure-based virtual screening (SBVS) is a cornerstone of computer-aided drug design, yet its success depends on selecting a combination of docking tools, scoring function (SF), and ranking strategies. MolDockLab addresses this challenge with an automated, data-driven framework that optimizes SBVS workflows for a pr...

Hamza Agha, Y. Ibrahim, Michael Backenköhler et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.