Skip to content

NextTopDocker: A Large-Scale Docking-Power Benchmark Reveals Limitations of Current End-to-End Machine-Learning Docking and the Strength of Hybrid Rescoring

Aug 2026 · Journal of Medicinal Chemistry · 0 citations · 52 references

TL;DR

“NextTopDocker” is presented, a large, up-to-date, open-access data set for docking-power assessment comprising 14,038 training and 5201 test entries across 3173 unique protein targets, constructed from the Protein Data Bank.

Abstract

Predicting three-dimensional binding orientations of drug-like molecules remains challenging in structure-based drug design. Despite methodological advances, docking performance is often assessed on small and outdated benchmarks. We present “NextTopDocker,” a large, up-to-date, open-access data set for docking-power assessment comprising 14,038 training and 5201 test entries across 3173 unique protein targets, constructed from the Protein Data Bank. Developed with open-source tools, it includes crystallographic structures, Smina-generated docking poses, and ligand-similarity-aware training subsets. We benchmarked four state-of-the-art machine-learning (ML) docking frameworks (DeepDock, Interformer, SurfDock, and Uni-Mol Docking v.2) against classical (Smina) and hybrid baselines (GNINA 1.3 and logistic regression using Smina and GNINA 1.3 scores). Interformer alone matched the docking power of logistic regression on Smina poses, while the others showed dependence on downstream physics-based correction. Most raw ML-generated poses displayed steric clashes and/or implausible geometries, highlighting the need for physics-informed constraints in autonomous docking. “NextTopDocker” is available at https://github.com/caominhtr/NextTopDocker and https://zenodo.org/records/17492994.

View source

Similar papers

Open access Jul 2026

Mavchen-1: A Conformational Ensemble Platform for Protein–Ligand Pose Prediction That Substantially Outperforms Static Structure Prediction in a Category-Stratified Benchmark

A category-stratified, statistically powered benchmark comparing pose prediction from receptor conformational ensembles against AlphaFold2, used as a matched static-structure baseline, across 29 protein–ligand systems spanning cryptic-pocket, induced-fit, water-mediated, and autoimmune-indication target classes is presented.

Ryan Varghese, Pooja Tiwary, Krishil Oswal · 0 citations
Open access Jul 2026

SurroDock: A Deep Learning Surrogate for Accelerated Pre-Docking Ligand Prioritization in Structure-Based Virtual Screening

Results indicate that 2D-based docking-score surrogate modeling can provide a reproducible and retrainable strategy for large-scale structure-based virtual screening by concentrating docking resources on a smaller, enriched subset of compounds.

Jongkeun Choi · 0 citations
Open access Aug 2026

Evaluating molecular docking for binding affinity predictions: a systematic analysis of key parameters and the utility of AlphaFold2 structures for the Schrödinger dataset

Molecular docking is one of the most established methods in computational drug discovery, due to its balance of speed and accuracy. However, the accuracy of docking results depends on a number of different parameters, and systematic reference data for comparisons to more advanced methods for binding affinity prediction are still scarce. This study assesses the impact of key parameters on the accuracy of binding free energy estimates from docking, using nine benchmark systems with 278 high-affinity ligands. Using the Molecular Operating Environment (MOE), we evaluated combinations of three receptor structures (two crystal structures, one AlphaFold2 model), two force fields, two scoring functions, two receptor flexibility settings, and two statistical evaluation schemes. The performance of the docking approaches is measured based on the squared Pearson’s correlation coefficient (R²), the root mean square error (RMSE) with respect to the experimental binding affinities, as well as the mean signed error (MSE) and Kendall’s tau for individual targets and the full dataset. The results show that the scoring function and the protein structure are the most important factors for binding affinity accuracy in rigid docking with the MOE software. Amber10:EHT and MMFF94x force fields had the same average Rmean2 value, but Amber10:EHT had a lower average RMSEmean. AlphaFold2 protein models yielded lower binding affinity accuracy and higher errors compared to experimental crystal structures, although induced fit docking improved results. Using the original benchmark, we also compared several docking programs. DOCK6 and MOE performed best, with mean R² values of about 0.49 and 0.40, respectively. The remaining docking programs did not outperform a molecular weight regression baseline. For a subset of four targets (CDK2, JNK1, P38, TYK2) evaluated in previous work, the performance of the optimized DOCK6 and MOE protocols produced correlation coefficients similar to those reported for certain MM/PBSA, FMO, and Boltz2 implementations evaluated on the same target subset. This raises questions about potential dataset biases, the structural preparation, or the implementation of those methods. Docking therefore should be considered as an important and computationally inexpensive reference baseline for binding affinity prediction.

Konstantinos Tornesakis, J. Essex, Paul A. Cox et al. · 0 citations
Open access Jul 2026

Rank-Resolved Multi-Engine Docking and Optuna-Optimized Re-Ranking with ProDock for Virtual Screening

False positives in virtual screening often arise when a single docking score or top-ranked pose is treated as sufficient evidence for binding. We extend the previously introduced ProDock software from a database-backed docking platform into a rank-resolved, multi-engine workflow for automated preparation, docking, pose analysis, and optimized re-ranking. The extended workflow combines local docking with GNINA and global docking with DiffDock with pose-level descriptors, namely binding-site occupancy, ligand localization, interaction-fingerprint similarity, and steric clash counts, together with Optuna -based threshold optimization. Across 43 DUDE-Z targets, the archived benchmark outputs reported higher enrichment values for CNN-based GNINA scores after optimization. CNNaffinity PR-AUC changed from 0.197 to 0.294 and LogAUC from 0.708 to 0.763, whereas empirical affinity ROC-AUC changed from 0.770 to 0.758. Structural investigation of re-docked actives showed that re-ranked poses were more native-like, with improved binding-site occupancy, reduced centroid displacement, and greater recovery of co-crystal interactions. The extension provides a reproducible framework for combining complementary docking engines with interpretable pose-level metrics before hit selection, thereby aiding the identification of true-positive candidates in virtual screening.

Lai Hoang Son Le, Thanh-An Pham, Ngoc Nguyen Tran et al. · 0 citations
Open access Aug 2026

Nesso-1: Accelerating Open-Source Binding Affinity Predictions

In this technical report, we introduce Nesso-1, a coarse-grained cofolding framework for binding- affinity prediction. Nesso-1 requires ∼ 1 second per prediction on a single GPU. This offers more than one order of magnitude speed-up over the leading open-source baseline, Boltz-2, which significantly expands the regions of chemical space that can be explored during high-throughput virtual screening. Importantly, Nesso-1 matches or surpasses the accuracy of Boltz-2 over the same benchmarks adopted in their study—which we show reflect in-distribution scenarios—as well as over more challenging out-of-distribution data encompassing the OpenBind affinity benchmark and 25 internal biochemical assays. Notably, Nesso-1 maintains robust predictive accuracy even on assays with extremely low similarity to the training data. Moreover, we highlight examples where Nesso-1 demonstrates meaningful selectivity, separating the binding affinities of identical compounds between on-targets and related off-targets. Nonetheless, zero-shot generalization to real- world medicinal chemistry remains an inherently challenging task; consequently, we acknowledge specific assays where the model’s performance is limited. We open-source Nesso-1: code and weights are available at https://github.com/recursionpharma/nesso

Nikhil Shenoy, David Errington, Emmanuel Bengio et al. · 0 citations