Skip to content
Open access

Physics‐based prediction of protein folding and unfolding rates: Examining the roles of fold topology and core packing

Jul 2026 · Protein Science · Vol 35 · 0 citations · 51 references
Medicine

Abstract

The (un)folding rates of natural proteins determine their native stability and functional homeostasis, making them important targets for protein engineering and design. From a prediction standpoint, the rates have been a long‐standing puzzle. We have known for decades that folding rates empirically correlate with properties of the native three dimensional (3D) structures and that both, folding and unfolding rates, scale with protein size. Whereas such rate correlations are too rough for being of practical use, no significant progress in prediction accuracy has occurred since then, despite many efforts even including machine learning approaches. Here, we retake on this challenge by expanding the simple one‐dimensional free energy surface (1D‐FES) model that originally led to demonstrate the size scaling of both rates, and a curated database with rates for 75 single‐domain proteins. We define the weighted sequence order (WSO) as a novel parameter that allows incorporating structural information into the 1D‐FES model explicitly. Via the WSO, we examine the role of global structural properties such as fold topology and core packing in defining the (un)folding rates within the context of a physics‐based model of protein folding. After introducing fold topology and packing at a coarse‐grained level, the model uses three floating parameters to predict the folding and unfolding rates within 6.5‐ and 10‐fold, respectively, resulting in ±6.5 kJ/mol accuracy in native stability, equivalent to the typical perturbation induced by one single‐point mutation. The net improvement over the 2‐parameter size‐only prediction is of 2.5‐fold. These new rate predictions are significantly closer to the threshold of usefulness for engineering and design. More importantly, this WSO‐modified 1D‐FES model can now directly accommodate atomistic, high‐resolution, force‐fields to further optimize the rate predictions, and/or to use rate information as a testbed for force‐field refinement. Finally, the WSO‐1D‐FES model could also serve as foundation for developing more complex models capable of dealing with multi‐domain proteins as well as with the evolutionary information cryptically encoded in natural protein sequences.

Read PDF

Similar papers

Open access Aug 2026

Blending physics-based and inverse folding models to disentangle variant effects on stability and function

Protein sequences are constrained not only by the need to fold into stable structures, but also by specific functional requirements imposed by natural selection. Yet predictions of how amino-acid changes affect proteins typically collapse these constraints into a single scalar score. Quantitatively separating these effects at scale remains an open challenge, with direct relevance spanning protein design to understanding the molecular mechanisms of disease. Inverse-folding (IF) models have emerged as fast, unsupervised predictors of folding energy changes (ΔΔG), but because they learn statistical correspondences between structure and sequence, they can conflate conservation driven by function with conservation driven by stability. Here, we show that blending IF models with a physics-based coarse-grained potential improves global correlation with experimental ΔΔG and, crucially, reduces IF model bias at functional sites. Applying the best-performing blend together with an evolutionary language model, we decompose each variant’s evolutionary cost into folding energy and dark energy, the latter capturing functional constraints beyond folding stability. With this decomposition, and without the need for supervision, we find that disease gain-of-function variants show a distinct functional signature from loss-of-function variants. In particular, we identify oncogenic drivers as largely preserving stability while exhibiting high dark energy, as opposed to tumor suppressors which are predominantly destabilized, paving the way to a mechanistic understanding of driver mutations in cancer. Together, these results provide a scalable framework for accurate ΔΔG prediction and mechanistic disentanglement of variant effects.

Ezequiel A. Galpern, Xavier Soler Sanchis, Charles W. J. Pugh et al. · 0 citations
Open access Jul 2026

RINAMI: Residue‐attributed interpretable neural network for predicting absolute folding free energy by merging structure and sequence information

Recent advances in de novo protein design have enabled the generation of diverse novel proteins. However, a fundamental challenge remains: even when an amino acid sequence is designed with the target structure as the most stable conformation, there is currently no reliable computational method for assessing whether the target structure is sufficiently stabilized relative to alternative conformations. While experimental realization of the intended fold requires the target structure to be thermodynamically favored by a large free‐energy gap, the absence of a quantitative measure of folding stability makes it difficult to distinguish reliable from unreliable designs. Here, we propose the Residue‐attributed Interpretable Neural network for predicting Absolute folding free energy by Merging structure and sequence Information (RINAMI), a machine learning model that predicts the absolute folding free energy (ΔG) of proteins from their three‐dimensional structures and amino acid sequences. RINAMI integrates structure‐ and sequence‐based representations derived from ProteinMPNN and Evolutionary Scale Modeling 2 (ESM2) using a multi‐head cross‐attention mechanism that contextualizes sequence‐derived signals within the structural environment. Benchmarking RINAMI on both natural and designed proteins from the Mega‐scale and Maxwell datasets shows that it outperforms the tested existing approaches, achieving higher correlations with experimental measurements and improved or comparable prediction errors. An ablation study supports the contribution of sequence–structure integration for predictive accuracy. In addition, RINAMI exhibits strong interpretability by capturing key physicochemical effects, including the destabilizing effect of buried hydrophilic residues, the stabilizing effect of buried hydrophobic residues, and the characteristics of cysteine. Together, these results establish RINAMI as an accurate and interpretable framework for ΔG prediction and provide a practical computational tool for evaluating and prioritizing protein designs prior to experimental testing.

Naoki Tomita, G. Chikenji · 0 citations
Open access

Experimental data-guided parameterization and validation of an AMBER protein force field

Folded or globular proteins adopt well-defined three-dimensional structures that correlate directly with function, forming the classical structure-function paradigm. Intrinsically Disordered Proteins (IDPs) and Regions (IDRs), encoded by roughly one-third of the eukaryotic genome, lack this fixed structure; their structural plasticity and major roles in diverse biological phenomena instead challenge this paradigm, making them a critical class of biomolecules studied through both experimental and computational approaches. Molecular dynamics (MD) simulations provide an atomistic understanding of IDP structure and dynamics, but their accuracy is limited by methodological assumptions and by the precision of the force field--the potential-energy function that approximates atomic interactions and governs how faithfully a simulation reflects real conformational behavior. Previous studies show that contemporary force fields fail to adequately capture amino acid residue specificity, limiting their applicability to IDPs. The central focus of this dissertation is therefore the development of an improved force field, Amber ff24EXP-GA, derived from its parent, Amber ff14SB, and its evaluation against Amber ff14SB and other contemporary force fields, such as CHARMM36m, in capturing the empirically determined conformational properties of unfolded systems: short peptides that serve as model systems for IDPs, and longer unfolded proteins. Chapter 3 addresses this gap by reporting a new force field, Amber ff24EXP-GA, derived from Amber ff14SB by optimizing backbone dihedral potentials for guest glycine and alanine residues in cationic GGG and GAG peptides, respectively, to best match guest-residue-specific spectroscopic data. Amber ff24EXP-GA outperforms Amber ff14SB for conformational ensembles of all 14 guest residues x (G, A, L, V, I, F, Y, D^p, E^p, R, C, N, S, T) in GxG peptides in water--the full set for which spectroscopic data exist--and outperforms CHARMM36m for at least 7 of them (G, A, V, F, C, T, E^p), showing greater amino acid specificity than both Amber ff14SB and CHARMM36m. It also reproduces experimental data on three-folded proteins and three longer IDPs well, while still outperforming Amber ff14SB on short unfolded peptides. Chapter 4 examines the effect of nearest-neighbor (NN) residues on conformational dynamics: extensive experimental evidence shows that the Flory hypothesis--which assumes neighboring residues behave independently--does not hold, as NN residues instead alter the conformational landscape of a given residue. Here, we evaluate CHARMM36m, Amber ff14SB, and Amber ff24EXP-GA for their ability to capture these NN effects on conformational dynamics of amino acid residues in short unfolded peptides in water. Amber ff24EXP-GA, whose reproduction of intrinsic conformational ensembles is significantly better than its parent's, also captures the NN effects better than Amber ff14SB. Despite lacking residue specificity in its intrinsic conformational ensembles, CHARMM36m--calibrated on global IDP properties--reproduces the NN effects on par with Amber ff24EXP-GA. These findings matter for the development of next-generation force fields that capture both residue-specific dynamics and global IDP properties. Chapter 5 examines the current scope, achievements, and limitations of various MD force-field parametrization strategies for modeling proteins, with particular emphasis on intrinsically disordered proteins.

Athul Suresh, B. Urbanc · 0 citations
Review Aug 2026

Local energetic frustration: Protein evolution, conformational dynamics and design in the age of AI.

Proteins operate under competing demands imposed by stability, dynamics, and function, all of which are shaped by evolution. Local energetic frustration provides a quantitative framework to describe how these competing requirements are distributed within the native states of proteins, identifying regions where interactions are optimized and others where energetic conflicts are retained to enable functional behavior. In recent years, the study of local frustration has expanded significantly, driven by the integration of large-scale structural datasets and advances in artificial intelligence methods. Comparative analyses have shown that frustration patterns encode evolutionary pressures across protein families, with minimally frustrated interactions stabilizing structural cores and highly frustrated regions often associated with catalysis, binding, conformational transitions as well as pathogenic phenotypes. At the same time, modern protein language models and structure prediction methods seem to implicitly capture the statistical and structural features underlying frustration, enabling its prediction directly from sequence or structure at proteome scale. These developments suggest that local energetic frustration may be interpreted as an emergent property of the evolutionary information learned by AI models. Here, we review recent advances in the analysis and prediction of local frustration and discuss how this framework could provide mechanistic insights into protein evolution, conformational dynamics, and design. We further argue that incorporating frustration into computational and experimental strategies will be essential to move beyond purely stability-driven approaches toward the rational engineering of functional proteins.

Franco L. Simonetti, Eli J. Draizen, Rocío Espada et al. · 0 citations