Machine learning (ML) models for molecular property prediction are increasingly deployed in drug discovery, yet their adoption in real-world scenarios requires an understanding of the conditions in which a model succeeds or fails. While standardized benchmarks are powerful instruments to measure and unlock progress in ML research, they should not be blindly treated as the end goal. Especially static and retrospective benchmarks, in which no true unknown test set is employed, limit our ability to robustly validate a model’s performance. Building on the collective expertise of a cross-industry consortium, we present a model validation framework consisting of five recommendations that would enable the community to move beyond aggregate metrics toward understanding where and why molecular property prediction models fail. We connect evaluation choices to real-world applications and case studies encountered in pharmaceutical research. The framework proposes splitting strategies that mimic realistic distribution shifts and expose common failure modes. We apply the recommended framework on a recently released dataset of absorption, distribution, metabolism, and excretion (ADME) properties. Across two complementary model algorithms, our case studies reveal four distinct failure modes (extrapolation, interpolation, representation, and evaluation) showing that model errors arise not only from distribution shift but also from limitations in molecular representations. Our results show that commonly used evaluation protocols can significantly overestimate performance and may not detect important model failure modes. All software and data are released via https://github.com/srijitseal/polaris.
Srijit Seal, Akshat Shirish Zalte, David Alencar Araripe et al.· bioRxiv· 0 citations
We present FLOWR.ROOT, an SE(3)-equivariant flow-matching foundation model that unifies pocket-aware 3D ligand generation with multi-endpoint binding affinity prediction (pIC50, pKi, pKd, pEC50) and pLDDT-based confidence estimation in a single backbone. One trained model supports de novo pocket-conditional generation, interaction- and pharmacophore-conditional sampling, scaffold hopping and elaboration, and fragment growing or replacement, enabled by a mixed isotropic–anisotropic prior placement strategy. Training proceeds in three stages: large-scale pre-training on billions of ligand conformations and millions of mixed-fidelity protein–ligand complexes, refinement on curated co-crystal data, and project-specific adaptation via parameter-efficient LoRA finetuning. Joint structure–affinity modelling enables inference-time importance-sampling guidance for single- and multi-objective design without external scoring functions. Case studies on kinase selectivity (CK2α/CLK3) and scaffold elaboration on TYK2, ERα, and BACE1 illustrate utility from hit identification through lead optimization. Structure-based generative modeling is rapidly reshaping drug discovery by enabling pocket-aware ligand design alongside predictive evaluation of binding properties. This manuscript introduces FLOWR.root, an SE(3)-equivariant flow-matching framework that jointly generates high-quality 3D ligands and predicts multi-endpoint binding affinities, demonstrating state-of-the-art performance, efficient domain adaptation, and practical impact across de novo design, scaffold elaboration, and lead optimization workflows.
Julian Cremer, Tuan Le, M. Ghahremanpour et al.· Nature Communications· 1 citation
Flowr.root achieves state-of-the-art performance in both unconditional 3D molecule and pocket-conditional ligand generation, producing geometrically realistic, low-strain structures with computational efficiency on established benchmark datasets.
Julian Cremer, Tuan Le, M. Ghahremanpour et al.· arXiv.org· 5 citations· ⚡1