DynaPPI, a dynamic protein dataset comprising molecular dynamics trajectories of protein complex formation from dissociated chains to the bound state, is presented as a pivotal resource to bridge the gap between static structural biology and the inherently temporal nature of dynamic molecular interactions.
Abstract
Diffusion models have been widely explored in protein backbone generation due to their powerful generation capabilities.However, in today's AI-driven biological research, predicting the structure of unknown multi-chain protein aggregates (called"complexes"in biology) remains an unsolved challenge.This is because existing static or dynamic protein datasets focus solely on static snapshots or single-entity trajectories, neglecting the dynamic process of multiple monomers forming complexes.To alleviate this dilemma, we present DynaPPI, a dynamic protein dataset comprising molecular dynamics (MD) trajectories of protein complex formation from dissociated chains to the bound state, as a pivotal resource to bridge the gap between static structural biology and the inherently temporal nature of dynamic molecular interactions.Benefiting from this dataset, diffusion models can explicitly learn the dynamic binding trajectories of known complexes and accurately predict the structures of unknown complexes based on their diverse generative properties, thereby further catalyzing AI-driven structural biology and protein interactomics.
This mini review traces the evolution of AI-driven methods in protein research, from early residue-contact prediction using coevolutionary information to transformative breakthroughs, the rise of protein language models (PLMs), and the emerging era of generative design and functional modeling.
Guodong Min, Huan Peng· Methods in molecular biology· 0 citations
Protein function emerges from dynamic conformational ensembles and transitions that are challenging to characterize experimentally and computationally. Recent advances in generative AI have created new opportunities for learning molecular thermodynamics, kinetics, and conformational evolution directly from simulation data, but progress is limited by the availability of large-scale datasets that combine rigorous sampling, complete phase-space information, and diverse physicochemical perturbations. Here, we present pHaseMD4AI, a molecular dynamics dataset that combines a globally equilibrated peptide branch with a protein-scale constant-pH molecular dynamics (CpHMD) branch spanning hundreds of soluble proteins. The peptide branch includes a complete set of canonical tripeptide and tetrapeptide systems together with post-translationally modified (PTM) and protonation-state datasets, providing synchronized atomic coordinates (R), velocities (V), forces (F), and Markov state model-based kinetic annotations. An accompanying web portal (https://isb.zju.edu.cn/md4ai/) enables users to browse, visualize, and download trajectories, annotations, and metadata. As an example application, we demonstrate a sequence-based model that can predict residue-level equilibrium dihedral distributions from sequence. pHaseMD4AI provides a resource for developing and benchmarking molecular machine learning methods while supporting broader studies of biomolecular dynamics under sequence, post-translational modification, and protonation-state perturbations.
Tiefeng Song, Yixin Guo, Jiahao He et al.· bioRxiv· 0 citations
A new physics-informed representation using Fourier transforms as an inductive bias for the multiscale temporal nature of protein dynamics, DynaMode is presented, achieving strong performance across a set of ensemble-based metrics.
H. Phipps, M. Cagiada, S. Villalba et al.· 0 citations
A pipeline reformulating kinase-substrate modeling as a Bayesian inference problem is presented and it is revealed that the interaction types and distances to the catalytic pocket significantly influence pathogenicity scores.
Jinyuan Hu, Shimian Li, Yue Xue et al.· Journal of Chemical Informat...· 0 citations
The utility of HA sites for suggesting candidate binding sites and the biological interpretability of PLM representations is explored, demonstrating the biological interpretability of PLM representations and offers a valuable method to prioritize functionally relevant protein residues for targeted biomedical research.
Sophia J. Pribus, Russ B. Altman, Gowri Nayar· bioRxiv· 0 citations
The state of a cell depends not only on protein abundance, but also on the biochemical and cellular activities of proteins, which are largely invisible to abundance profiling alone. Here, we introduce a multi-omics framework that infers context-specific protein activities from transcriptomic, phosphoproteomic, and protein correlation-based protein-protein interaction data, integrating modality-specific algorithms via network diffusion. Applying it to a panel of phenotypically diverse HeLa cell lines, whose genetic drift provides a natural perturbation system, we make three findings. First, physical separation of monomeric and assembled protein fractions by protein correlation profiling provides direct evidence that complex assembly buffers variation in gene copy number and transcription, a mechanism previously only inferred from bulk measurements. Second, using Let7 perturbation data, CRISPR gene dependency scores, and subcellular localization, we orthogonally validate that inferred protein activities capture functional regulation linked to cellular phenotypes inaccessible from abundance data alone. Third, differential analysis of context-specific activity profiles identifies molecular mechanisms underlying phenotypic divergence, including a WIPF1/WIPF2--Arp2/3 axis governing invadopodium formation and infection susceptibility, and an immunoproteasome switch linked to immune adaptation.
George A. Rosenberger, Peng Xue, Isabell Bludau et al.· Molecular Systems Biology· 0 citations