Jul 2026· ACS Central Science· Vol 12, pp. 1178 - 1194· 0 citations· 59 references
Medicine
TL;DR
A similarity-based prioritization approach, ChemSTEP, is introduced, and the effective size of a library treated by any prioritization algorithm, Neff is defined, which suggests trillion-molecule libraries might be in reach with this approach.
Abstract
While make-on-demand libraries now span trillions of molecules, full library docking struggles beyond a few billion, motivating prioritization that recovers top-scoring compounds while evaluating only a fraction of a library. Here we introduce a similarity-based prioritization approach, ChemSTEP, and define the effective size of a library treated by any prioritization algorithm, Neff. ChemSTEP docks a representative seed set, selects diverse high-scoring “beacons”, and iteratively traverses the library through cycles of beacon selection, similarity search, and docking. Retrospectively on eight targets, ChemSTEP recovered over 75% of high-scoring compounds while docking less than 5% of a library. We then tested ChemSTEP prospectively against AmpC β-lactamase using a 13.2 billion molecule library. Because AmpC recognizes negatively charged inhibitors, we explicitly docked all 360 million library anions, synthesizing and testing 241 high-ranking ones in parallel to the ChemSTEP 13.2B run. Compared with previous docking of 99 million and 1.7 billion molecules against AmpC, the 13.2 billion library had higher hit-rates (2% vs 25% vs 37%, respectively) and found more potent compounds. Meanwhile, ChemSTEP retrieved 80% of the 241 high-ranking compounds within the first 0.5% docked. Trillion-molecule libraries might be in reach with this approach.
As chemical libraries expand into the trillions of molecules, Virtual SYNthon Hierarchical Enumeration Screening (V-SYNTHES) has emerged as a leading strategy for making gigascale virtual screening computationally tractable. In V-SYNTHES, a Minimal Enumeration Library (MEL) of chemical fragments is docked against a tar...
Wen-Jin Liu, Yong-Chan Hong, Thomas Ku et al.· bioRxiv· 0 citations
Virtual screening of chemical libraries has long been a cornerstone for identifying candidate bioactive compounds. Recently, many tools have been developed to make virtual screening more accessible. Among them, the jamdock-suite provides an automated and user-friendly workflow that encompasses the entire process from l...
Shahariar Emon, M. Haque, Iftekhar Alam et al.· STAR Protocols· 0 citations
The rapid expansion of large chemical libraries has created a need for virtual screening workflows that are both efficient and accurate. Active learning (AL) offers a scalable strategy by iteratively training surrogate models to prioritize promising compounds and reduce the number of required docking calculations. Howe...
Joseph Chung, Aashish Bhatt, Jacob Ede Levine et al.· Journal of Computer-Aided Mo...· 0 citations
Structure-based virtual screening (SBVS) is a cornerstone of computer-aided drug design, yet its success depends on selecting a combination of docking tools, scoring function (SF), and ranking strategies. MolDockLab addresses this challenge with an automated, data-driven framework that optimizes SBVS workflows for a pr...
Hamza Agha, Y. Ibrahim, Michael Backenköhler et al.· npj Drug Discovery· 0 citations
PandaDock’s empirical scoring function ranks 8th of 25 methods evaluated, ahead of every AutoDock Vina and Vinardo configuration tested, while the GNN scores below Vina, consistent with the within-target ceiling identified on SAIR.
P. Panda· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.