Skip to content
Open access

FLOWR.ROOT – A flow matching-based foundation model for joint multi-purpose structure-aware 3D ligand generation and affinity prediction

Jul 2026 · Nature Communications · Vol 17 · 1 citation · 91 references
Medicine

Abstract

We present FLOWR.ROOT, an SE(3)-equivariant flow-matching foundation model that unifies pocket-aware 3D ligand generation with multi-endpoint binding affinity prediction (pIC50, pKi, pKd, pEC50) and pLDDT-based confidence estimation in a single backbone. One trained model supports de novo pocket-conditional generation, interaction- and pharmacophore-conditional sampling, scaffold hopping and elaboration, and fragment growing or replacement, enabled by a mixed isotropic–anisotropic prior placement strategy. Training proceeds in three stages: large-scale pre-training on billions of ligand conformations and millions of mixed-fidelity protein–ligand complexes, refinement on curated co-crystal data, and project-specific adaptation via parameter-efficient LoRA finetuning. Joint structure–affinity modelling enables inference-time importance-sampling guidance for single- and multi-objective design without external scoring functions. Case studies on kinase selectivity (CK2α/CLK3) and scaffold elaboration on TYK2, ERα, and BACE1 illustrate utility from hit identification through lead optimization. Structure-based generative modeling is rapidly reshaping drug discovery by enabling pocket-aware ligand design alongside predictive evaluation of binding properties. This manuscript introduces FLOWR.root, an SE(3)-equivariant flow-matching framework that jointly generates high-quality 3D ligands and predicts multi-endpoint binding affinities, demonstrating state-of-the-art performance, efficient domain adaptation, and practical impact across de novo design, scaffold elaboration, and lead optimization workflows.

Read PDF

Similar papers

Oct 2025

Flowr.root – A flow matching based foundation model for joint multi-purpose structure-aware 3D ligand generation and affinity prediction

Flowr.root achieves state-of-the-art performance in both unconditional 3D molecule and pocket-conditional ligand generation, producing geometrically realistic, low-strain structures with computational efficiency on established benchmark datasets.

Julian Cremer, Tuan Le, M. Ghahremanpour et al. · 5 citations · ⚡1
Preprint Jul 2026

Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints

A clear pattern is revealed in LLM spatial capabilities: while they still lag behind state-of-the-art approaches, they are promising and can handle multiple spatial constraints simultaneously, enabling scaling to heterogeneous setups.

Thomas MacDougall, Maksim Kuznetsov, Roman Schutski et al. · 0 citations
Open access Aug 2026

Structure-agnostic protein–ligand binding affinity prediction via hierarchical representation alignment

Abstract Motivation To enable real-world protein-ligand affinity prediction, not only out-of-distribution generalization but also robustness to variable structural availability and quality should be considered in model design. Results We present AlignNet, a hierarchical representation alignment framework that mitigates intra- and inter-molecular heterogeneity to learn robust protein-ligand embeddings for generalizable affinity prediction, even from sequence-level inputs. Its intra-molecular module projects unimodal and multimodal features into a unified space, aligning augmented multimodal views for feature fusion and unimodal with multimodal embeddings to distill multimodal priors for structure-agnostic inference. Its inter-molecular module aligns protein and ligand embeddings for cross-molecular integration. Extensive experiments show that AlignNet (i) achieves highly competitive performance, with up to a 20.4% gain in SCC on the challenging LBA 30% split under sequence-only settings, suggesting improved out-of-distribution generalization; and (ii) learns well-separated affinity-related clusters, supporting reliable structure-independent prediction. Availability and implementation AlignNet is available at https://github.com/altriavin/AlignNet.

Xiaowen Hu, Hongyi Huang, Hao Sun et al. · 0 citations
Open access Jul 2026

A Preparation-Free Mixture-of-Experts Framework for Protein-Ligand Affinity Prediction

The resulting model, HydrAffinity, is an interaction-free, dynamic sparse model that uses pre-trained encoders and MoE for parameter-efficient learning and outperforms all interaction-free methods and matches state-of-the-art interaction-based methods on CASF-2016.

Huiming Bao, Shouliang Dong · 0 citations
Jun 2025

READ: A Retrieval-Alignment Diffusion Framework for Structure-based Drug Design.

Structure-based drug design (SBDD) models are central to modern pharmaceutical research, enabling the rational exploration of protein-ligand interactions at atomic resolution. However, most existing approaches frame molecular generation as an isolated optimization or a one-to-one matching task, overlooking the shared binding patterns and intrinsic similarities among protein-ligand complexes. This fragmented perspective constrains their ability to capture the fundamental principles governing molecular recognition and binding specificity. Moreover, the limited availability of high-quality experimental data further hampers model generalization and real-world applicability. To address these challenges, we present READ, a retrieval-alignment molecular generation framework that conditions the generative process on small molecules targeting homologous proteins. Retrieved ligands are aligned with a diffusion model across multiple representational spaces and integrated as conditional guidance throughout successive stages of generation. Under a standardized docking-based evaluation protocol, READ achieves consistently strong performance against state-of-the-art SBDD methods. More importantly, it introduces a retrieval-alignment paradigm for structure-based molecular generation, offering a practical framework for early-stage computational hit generation while leaving prospective experimental validation as future work.

Dong Xu, Zhangfan Yang, Junchuang Cai et al. · 1 citation
Open access Aug 2026

LEN-Seek: Fast and scalable ligand binding-site similarity search in the latent space of an SE(3)-invariant graph VAE

Motivation Ligand binding-site similarity search is a crucial step in drug discovery that reduces the conformational search space for docking and other downstream tasks by comparing a target protein against experimentally identified binding sites. Existing methods rely on either direct structural alignment or lossy compression of structural information, producing a trade-off between scalability and precision. Results We propose LEN-Seek, a ligand binding-site search method based on a graph neural network (GNN)-driven variational autoencoder (VAE) that encodes the 3D structural and physicochemical context of a binding site into a probabilistic latent space, enabling similarity search within a low-dimensional vector space. A binding site is modeled as a graph of amino acid residues, with node features adopted from the protein language model, Ankh, and edges encoded as SE(3)-invariant (roto-translational invariant) geometric relationships, thereby avoiding expensive data augmentation or SE(3)-equivariant models. Compared to ProBiS, the purely geometric graph-clique based method, LEN-Seek successfully retrieves a substantial portion of similar binding sites with a roughly 3,400-fold lower per-comparison cost, demonstrating its potential as a scalable approach to template-based ligand binding-site search in large-scale protein structure databases. Supplementary information Supplementary data are available at Bioinformatics online.

Kyunghwan Yeo, Dongwoo Kim, Jaemin Sim et al. · 0 citations