Skip to content
Preprint

Spectral Diffusion for Protein Dynamics

Jul 2026 · 0 citations · 50 references
Biology

TL;DR

A new physics-informed representation using Fourier transforms as an inductive bias for the multiscale temporal nature of protein dynamics, DynaMode is presented, achieving strong performance across a set of ensemble-based metrics.

Abstract

Generative models present a promising alternative to expensive molecular dynamics for computationally querying protein dynamics, yet many existing approaches treat ensembles as unordered snapshots rather than temporally coherent trajectories, or scale poorly with protein size. We present a new physics-informed representation using Fourier transforms as an inductive bias for the multiscale temporal nature of protein dynamics. Diffusion in the spectral domain allows for disentangling of dynamics into slow conformational modes and fast atomic jitter, enabling rapid and improved prediction of dynamics across a range of temperatures. This is facilitated by denoising of structure and temperature conditioned spectral volumes where the low frequencies directly encode per-residue flexibility. Trained on the mdCATH dataset, we evaluate our model, DynaMode, on a held-out test set achieving strong performance across a set of ensemble-based metrics including a Root Mean Squared Fluctuation (RMSF) pearson $r$ of $0.844$. Code is available at https://github.com/HPuntu/DynaMode.

View source

Similar papers

Preprint Jul 2026

DyneTrion: A Spatio-temporally Coherent Generative Emulator for Protein Dynamics Across Timescales

Proteins function through coordinated motion across multiple spatial and temporal scales, underpinning processes such as ligand binding, allostery, and catalysis. However, accessing long-timescale conformational change through molecular dynamics (MD) simulations remains prohibitively expensive for systematic exploration across diverse systems. Here, we present DyneTrion, a generative protein dynamics emulator that jointly enforces geometric symmetry, structural consistency and temporal coherence within a single framework. DyneTrion uses a tri-attention architecture that integrates invariant point attention (IPA) for SE(3)-robust geometric updates, spatial attention anchored to a reference conformation to preserve structural integrity, and temporal attention to model correlated evolution across time frames. Across 100-ns MD trajectory simulation benchmarks, DyneTrion reproduces MD-derived flexibility, ensemble distributions and interaction observables while maintaining stereochemical validity during extrapolation. To evaluate long time-scale generalization, we introduce dynamicPDB, a dataset of over 10,000 proteins with up to 1-$\mu$s all-atom trajectories at 10-ps resolution and accompanying physical annotations. On microsecond trajectories, DyneTrion preserves free-energy landscapes and metastable-state populations, and it supports large conformational propagation in apo-to-holo transitions and fast folders. Together, DyneTrion provides a scalable path from static structure prediction toward time-resolved, ensemble-faithful protein modeling. The code is publicly available at https://github.com/fudan-generative-vision/DyneTrion

Kaihui Cheng, Zhiqiang Cai, Peng Tu et al. · 0 citations
Open access Jul 2026

pHaseMD4AI: Phase-Space Dynamics Dataset with Chemical and pH Perturbations for Physically and Kinetically Consistent Biomolecular AI

Protein function emerges from dynamic conformational ensembles and transitions that are challenging to characterize experimentally and computationally. Recent advances in generative AI have created new opportunities for learning molecular thermodynamics, kinetics, and conformational evolution directly from simulation data, but progress is limited by the availability of large-scale datasets that combine rigorous sampling, complete phase-space information, and diverse physicochemical perturbations. Here, we present pHaseMD4AI, a molecular dynamics dataset that combines a globally equilibrated peptide branch with a protein-scale constant-pH molecular dynamics (CpHMD) branch spanning hundreds of soluble proteins. The peptide branch includes a complete set of canonical tripeptide and tetrapeptide systems together with post-translationally modified (PTM) and protonation-state datasets, providing synchronized atomic coordinates (R), velocities (V), forces (F), and Markov state model-based kinetic annotations. An accompanying web portal (https://isb.zju.edu.cn/md4ai/) enables users to browse, visualize, and download trajectories, annotations, and metadata. As an example application, we demonstrate a sequence-based model that can predict residue-level equilibrium dihedral distributions from sequence. pHaseMD4AI provides a resource for developing and benchmarking molecular machine learning methods while supporting broader studies of biomolecular dynamics under sequence, post-translational modification, and protonation-state perturbations.

Tiefeng Song, Yixin Guo, Jiahao He et al. · 0 citations
Open access Jul 2026

UniFlow: Unifying protein conformational ensemble generation and machine-learned force fields with a scalable normalizing Flow

UniFlow is introduced, the first scalable generative model that unifies protein ensemble generation and machine-learned coarse-grained force fields for molecular dynamics simulation within a single framework, and paves the way for a unified class of models that bridges generative ensemble modeling with physics-based molecular simulation.

Yikai Liu, Ming Chen, Guang Lin · 0 citations
Preprint Jul 2026

AquaGen: Scaling generative models to molecular dynamics precision on thousands of atoms

We present AquaGen, the first all-atom, explicit solvent, periodic-boundary-condition-aware generative model that produces molecular configurations from the Boltzmann distribution at a fraction of the cost of molecular dynamics (MD). This is in contrast with existing generative models that remove degrees of freedom by operating on coarse-grained, vacuum, or implicit solvent systems. Operating at this resolution allows for post-processing through force field energy evaluations and MD simulations, and enables the prediction of relevant properties in a gray-box manner (as ensemble averages of potential energy evaluations over generated samples). We demonstrate the utility of this paradigm on absolute hydration free energy (AHFE), producing estimates 4-10x faster and with comparable accuracy to standard GPU-based MD. By generating uncorrelated samples from alchemical Boltzmann distributions, we create more accurate, interpretable, and refinable ensemble predictions with calibrated uncertainty estimates, unlike regression methods which are entirely black-box predictors. Our approach also yields predictable benefits from increasing train- and test-time compute, realized by scaling model size and generating more samples, respectively. We believe that this approach demonstrates the utility of high-resolution ensemble generation for free energy estimation, with future potential to replace MD in tasks such as the prediction of lipophilicity, membrane permeability, or absolute binding free energy (ABFE) -- whose grounding and interpretability may be critical for the development of new drugs and materials.

Emmanuel Bengio, Sanjeev Raja, Y. Pang et al. · 1 citation
Open access Aug 2026

SPINDLE: Unlocking protein dynamics from single-field NMR relaxation data using a deep learning ensemble

A protein’s function is derived from its three-dimensional structure and the motions of the atoms about that structure. The detailed characterization of both macromolecular structure and dynamics provides an opportunity for understanding enzyme catalysis, ligand binding, and allostery, along with providing insights into how the function changes upon mutation or post-translational modification. Among the various methods for characterizing biomolecular motions, nuclear magnetic resonance (NMR) spin relaxation methods are a standard for determining nanosecond global tumbling times along with the amplitude and timescale of faster local motions. Within the model-free formalism, various mathematical models are used to extract dynamic parameters. Unfortunately, as the number of fitted parameters increases within these models, they become mathematically underdetermined for standard NMR relaxation data collected at a single magnetic field, necessitating multi-field datasets. Here, we present SPINDLE, an ensemble of deep neural networks trained on a large synthetic set of NMR relaxation data. Unlike traditional least-squares fitting, SPINDLE predicts both fast and slow timescale dynamics parameters from a single set (i.e., collected at a single magnetic field) of three relaxation datasets using the ensemble for error estimation. We demonstrate a strong correlation to ground truth dynamics parameters on synthetic benchmarks, with more precision than traditional fitting techniques, and precisely reproduce experimental dynamics parameters for ∼50 proteins with relaxation data in the Biological Magnetic Resonance Data Bank. We also leverage the architecture of the deep neural network to show how the model emphasizes rigid residues for the prediction of global correlation times. This strategy may be useful in the future for elucidating correlated networks of dynamic residues from multiple relaxation datasets.

Olivia E. Krise, Michael P. Latham · 0 citations
Open access Aug 2026

Inferring protein ensembles directly from NOESY spectra

Solution NMR spectroscopy provides atomistic measurements of proteins in a native-like biophysical state. Because these measurements are ensemble averages, it also has the potential to report on conformational diversity. However, conventional NMR structure determination typically converts experimental observables into restraints for molecular dynamics, which encode information on the mean structure but do not retain information on the underlying conformational distribution. Ensemble selection has long been proposed as an alternative, whereby experimental observables are compared directly with candidate conformers generated independently of the measurements. This allows population distributions to be inferred from the data. However, few such methods have incorporated NOESY - the richest source of structural information in protein NMR - data, due to challenges in the quantitative comparison of experimental and back-calculated spectra. To address this challenge, we previously introduced the CoMAND method, demonstrating that quantitative agreement is practical for NOESY spectra with bespoke heteronuclear editing schemes. Here we extend this approach into a framework for direct inference of protein ensembles within a flexible ensemble-selection architecture incorporating multiple classes of NMR observables. We introduce a quantitative scoring framework for comparing experimental and back-calculated observables and combine it with regularized ensemble selection and Monte Carlo simulated annealing. Integration with the OpenMM molecular dynamics engine allows conformational pools to be generated using established molecular simulation methods. Applied to human ubiquitin, the resulting ensemble provides simultaneous agreement with NOESY, residual dipolar coupling and scalar coupling data while retaining conformational diversity supported by experiment.

Murray Coles · 0 citations