We present AquaGen, the first all-atom, explicit solvent, periodic-boundary-condition-aware generative model that produces molecular configurations from the Boltzmann distribution at a fraction of the cost of molecular dynamics (MD). This is in contrast with existing generative models that remove degrees of freedom by operating on coarse-grained, vacuum, or implicit solvent systems. Operating at this resolution allows for post-processing through force field energy evaluations and MD simulations, and enables the prediction of relevant properties in a gray-box manner (as ensemble averages of potential energy evaluations over generated samples). We demonstrate the utility of this paradigm on absolute hydration free energy (AHFE), producing estimates 4-10x faster and with comparable accuracy to standard GPU-based MD. By generating uncorrelated samples from alchemical Boltzmann distributions, we create more accurate, interpretable, and refinable ensemble predictions with calibrated uncertainty estimates, unlike regression methods which are entirely black-box predictors. Our approach also yields predictable benefits from increasing train- and test-time compute, realized by scaling model size and generating more samples, respectively. We believe that this approach demonstrates the utility of high-resolution ensemble generation for free energy estimation, with future potential to replace MD in tasks such as the prediction of lipophilicity, membrane permeability, or absolute binding free energy (ABFE) -- whose grounding and interpretability may be critical for the development of new drugs and materials.
Emmanuel Bengio, Sanjeev Raja, Y. Pang et al.· 1 citation
In this technical report, we introduce Nesso-1, a coarse-grained cofolding framework for binding- affinity prediction. Nesso-1 requires ∼ 1 second per prediction on a single GPU. This offers more than one order of magnitude speed-up over the leading open-source baseline, Boltz-2, which significantly expands the regions of chemical space that can be explored during high-throughput virtual screening. Importantly, Nesso-1 matches or surpasses the accuracy of Boltz-2 over the same benchmarks adopted in their study—which we show reflect in-distribution scenarios—as well as over more challenging out-of-distribution data encompassing the OpenBind affinity benchmark and 25 internal biochemical assays. Notably, Nesso-1 maintains robust predictive accuracy even on assays with extremely low similarity to the training data. Moreover, we highlight examples where Nesso-1 demonstrates meaningful selectivity, separating the binding affinities of identical compounds between on-targets and related off-targets. Nonetheless, zero-shot generalization to real- world medicinal chemistry remains an inherently challenging task; consequently, we acknowledge specific assays where the model’s performance is limited. We open-source Nesso-1: code and weights are available at https://github.com/recursionpharma/nesso
Nikhil Shenoy, David Errington, Emmanuel Bengio et al.· bioRxiv· 0 citations