High-fidelity quantum chemical (QM) data sets that jointly resolve reaction thermochemistry, kinetics, and solvation at scale remain scarce, especially for radical chemistry. We introduce QuantumPioneer, an open-access reaction-centered QM database and workflow for small organic molecules, focused on peroxyl-mediated hydrogen atom transfer (HAT) and the corresponding homolytic bond dissociation reactions. QuantumPioneer contains 348,258 species (2–21 heavy atoms), 167,237 validated HAT transition states (TS) with corresponding reaction energies and homolytic bond dissociation energies (BDEs), and over 100 million COSMO-RS solvation free energies(ΔGsolv*) and enthalpies (ΔHsolv*) across 295 solvents. The workflow uses ωB97X-D/def2-SVP geometries, DLPNO–CCSD(T)-F12d/cc-pVTZ-F12 single-point energies, empirical thermochemical corrections, transition-state theory, and COSMO-RS BP-TZVPD-FINE solvation in a single high-throughput pipeline. Our benchmarks show reliable accuracy, with mean absolute errors (MAEs) compared to experimental and high-level QM reference data of 0.82 kcal/mol for gas-phase enthalpies of formation, 1.60 kcal/mol for C–H BDEs, 1.45 kcal/mol for HAT barriers, and 0.57 kcal/mol for ΔGsolv* values. We demonstrate two predictive applications. First, we show that combining BDE and HAT-barrier models identifies experimentally observed oxidative degradation sites in drug-like molecules with a 91% top-5 hit rate and 82% site-level recall. Second, we show that a QM-parametrized Abraham model enables rapid solvation energy estimates at near-COSMO-RS accuracy within its training domain, reproducing computed ΔGsolv* and ΔHsolv* values with MAEs of 0.16 and 0.18 kcal/mol, respectively, though performance on experimental ΔGsolv*values for unseen solutes was worse, with an MAE of 1.32 kcal/mol. This work provides a scalable template for other reaction families, unifying equilibrium species, validated TS, thermochemistry, kinetics, and solvation into one workflow.
Haoyang Wu, Jonathan W. Zheng, Hao‐Wei Pang et al.· Journal of the American Chem...· 1 citation
Machine learning (ML) models for molecular property prediction are increasingly deployed in drug discovery, yet their adoption in real-world scenarios requires an understanding of the conditions in which a model succeeds or fails. While standardized benchmarks are powerful instruments to measure and unlock progress in ML research, they should not be blindly treated as the end goal. Especially static and retrospective benchmarks, in which no true unknown test set is employed, limit our ability to robustly validate a model’s performance. Building on the collective expertise of a cross-industry consortium, we present a model validation framework consisting of five recommendations that would enable the community to move beyond aggregate metrics toward understanding where and why molecular property prediction models fail. We connect evaluation choices to real-world applications and case studies encountered in pharmaceutical research. The framework proposes splitting strategies that mimic realistic distribution shifts and expose common failure modes. We apply the recommended framework on a recently released dataset of absorption, distribution, metabolism, and excretion (ADME) properties. Across two complementary model algorithms, our case studies reveal four distinct failure modes (extrapolation, interpolation, representation, and evaluation) showing that model errors arise not only from distribution shift but also from limitations in molecular representations. Our results show that commonly used evaluation protocols can significantly overestimate performance and may not detect important model failure modes. All software and data are released via https://github.com/srijitseal/polaris.
Srijit Seal, Akshat Shirish Zalte, David Alencar Araripe et al.· bioRxiv· 0 citations
This work proposes pretraining on low-noise, calculable molecular descriptors via supervised learning to obtain rich, highly transferable molecular representations and demonstrates this strategy with CheMeleon, a O(10M) parameter foundation model that enables directed message-passing neural networks to finally exceed the performance of classical methods in the low-data regime.
Jackson W. Burns, Akshat Shirish Zalte, C. Abreu et al.· Journal of Chemical Informat...· 0 citations