Skip to content
Open access

Hierarchical Discrete Representations for Coarse-to-Fine Protein Conformation Generation.

Aug 2026 · Bioinformatics · 0 citations
Medicine

TL;DR

This work proposes a novel approach that learns hierarchical discrete representations of protein structures using vector quantization, and outperforms state-of-the-art models such as ESMDiff across challenging benchmark datasets, including BPTI MD trajectories and conformational-changing pairs.

Abstract

MOTIVATION Protein conformation generation remains a fundamental challenge in structural biology and machine learning. Proteins are highly dynamic macromolecules, and their functions are governed not only by static 3D structures but also by their conformational flexibility. Therefore, generating a diverse ensemble of physically plausible structures is crucial for modeling conformational heterogeneity, characterizing intrinsically disordered proteins (IDPs), and supporting structure-based drug discovery. However, existing generative models often operate in continuous coordinate space. This approach makes it difficult to impose structural priors or handle complex, long-range dependencies, frequently leading to a significant trade-off where models must sacrifice structural validity to ensure conformational diversity. RESULTS To overcome these limitations, we propose a novel approach that learns hierarchical discrete representations of protein structures using vector quantization. Inspired by the success of VQ-VAE models in vision, we discretize residue-level structural contexts into learnable codebooks at multiple levels of granularity. Our framework first constructs a coarse-grained scaffold that captures global topology and secondary structures, and subsequently conditions on this scaffold to progressively refine the fine-grained local geometry. This coarse-to-fine generation mechanism enables both efficient sampling and highly accurate reconstruction. Through extensive experiments, we demonstrate that our method outperforms state-of-the-art models such as ESMDiff across challenging benchmark datasets, including BPTI MD trajectories and conformational-changing pairs (Apo/holo and fold-switch). Our model mitigates the trade-off between conformational diversity and structural validity, generating realistic ensembles while providing interpretable and scalable representations of protein geometry. Ultimately, this work offers a powerful new paradigm for flexible biomolecular modeling. AVAILABILITY The code is available at https://github.com/vanha9/hierarchical_conformation_generation. Archived at https://zenodo.org/records/21817083.

Read PDF

Similar papers

Open access Jul 2026

Density-driven support fields for topological stability in protein structures

We model protein structural stability as a continuous scalar quantity defined over molecular geometry, referred to as the support field. Instead of treating stability as a discrete residue annotation or an empirical score, this representation characterizes protein folds through the combined effects of geometric organization, topological persistence, and local density. Based on this idea, we introduce Support Field Neural Representation Learning (SF-NRL), a topology-guided approach that integrates persistent homology(PH), spatial density estimation, and geometric deep learning to infer residue-wise support directly from protein structures. Persistent topological features are incorporated as structural constraints that modulate local support values across the fold, enabling a continuous description of structural reliability. Across diverse protein families, the inferred support field shows consistent agreement with independent indicators of structural stability and highlights low-support regions associated with conformational flexibility and weak structural integration. By embedding protein structures into a continuous stability landscape, SF-NRL provides an interpretable representation that complements structure prediction models and facilitates systematic identification of structural cores, flexible regions, and functionally relevant motifs. These results demonstrate that topology-informed field representations offer a generalizable and practically useful approach for analyzing protein stability and fold organization.

Jianshi Wang, Yukio Ohsawa · 0 citations
Review Open access Aug 2026

Diffusion-Based Protein Structure Design: Geometric Modelling, Validation Strategies, and Thermodynamic Challenges

This review focuses on coordinate- and residue-frame-based diffusion approaches for generating protein structures, paying particular attention to geometric equivariance, conditioning strategies, all-atom modelling and interaction-aware design.

Wen-Ran Li, Xavier F. Cadet, David Medina-Ortiz et al. · 0 citations
Open access Jul 2026

UniFlow: Unifying protein conformational ensemble generation and machine-learned force fields with a scalable normalizing Flow

UniFlow is introduced, the first scalable generative model that unifies protein ensemble generation and machine-learned coarse-grained force fields for molecular dynamics simulation within a single framework, and paves the way for a unified class of models that bridges generative ensemble modeling with physics-based molecular simulation.

Yikai Liu, Ming Chen, Guang Lin · 0 citations
Preprint Aug 2026

PHASE: encoding global protein ensembles with local Hamiltonians and all-atom backmapping

Protein function is governed by conformational ensembles, which can be viewed as high-dimensional probability distributions over molecular conformations. Yet the statistical organization of these distributions is often represented only implicitly, either through collections of simulation trajectories or within high-capacity generative models. Here, we introduce PHASE (Protein Hamiltonians for Sampling of Ensembles), a system-specific framework that converts atomistic conformational ensembles into an explicit and interpretable statistical model. Applied to ten conformational ensembles derived from approximately 37$\mu$s of atomistic simulations of the adenosine A2A receptor, Hamiltonians containing only local residue couplings within 6$\mathring{A}$ reproduce residue-wise and pairwise microstate statistics, including correlations between residues that are not directly coupled in the model. Moreover, independently fitted inactive and active reference Hamiltonians define an endpoint preference coordinate that organizes newly sampled ligand-, effector- and conformation-dependent ensembles along the A2A activation landscape without receiving these biochemical labels as model inputs. Finally, a cluster-conditioned all-atom reconstruction model preserves the prescribed residue microstate patterns of newly sampled configurations, closing the coarse-graining-sampling-backmapping cycle. The resulting discrete representation additionally admits direct QUBO encoding, enabling classical annealing and providing a route toward future quantum-annealing implementations. PHASE therefore provides a protein-general procedure for constructing compact, interpretable and atomistically realizable statistical models of protein conformational ensembles.

Daniele Angioletti, Marco S. Nobile, Matteo Carli et al. · 0 citations
Open access Aug 2026

Pi-Ensemble: Sequence-guided generation of interpolated protein conformational ensembles

Proteins are critical biomolecular machines that populate ensembles of interconverting conformations. Many biological processes depend on transitions between metastable states. Although molecular dynamics (MD) simulations provide a physically grounded route to characterize these motions, routine sampling of large-scale conformational transitions remains computationally demanding. Recent advances in protein structure prediction have created new opportunities for ensemble generation, but many existing approaches require noising inputs, task-specific training, supervised fitting on extensive MD data, or experimentally-informed restraints. Here, we introduce Pi-Ensemble (Predicting Interpolated Ensemble), a sequence-guided framework for generating protein conformational ensembles interpolating between two structural anchor states. Unlike previous methods, Pi-Ensemble alternately leverages inverse-folding and structure-prediction models to propose intermediate conformations between known protein states, generating diverse ensembles without additional training. We evaluate Pi-Ensemble across diverse protein systems, including enzymes, transporters, receptors, and benchmark cases with reference MD simulations or experimental Double Electron-Electron Resonance (DEER) data. Pi-Ensemble recovers physically plausible intermediate conformations, captures transition pathways observed in large-scale MD simulations, and generates structures consistent with experimental distance distributions. Furthermore, Pi-Ensemble-generated conformations provide effective starting seeds for parallel MD simulations, improving conformational exploration and accelerating convergence relative to simulations initiated only from endpoint structures. These results establish sequence-guided structural interpolation as a practical strategy for probing protein conformational landscapes. By generating diverse and physically reasonable conformational proposals without long-timescale MD or model retraining, Pi-Ensemble provides an extensible framework for studying protein flexibility, guiding adaptive sampling, and accelerating mechanistic investigations of protein function.

Hassan Nadeem, D. Kleiman, Yuming Zhou et al. · 0 citations