By smoothing the Evoformer's weight tensors with a Gaussian convolution and scaling the result, it is shown that the trained model produces physically structured conformational landscapes, appearing to encode structural constraints that extend beyond what unperturbed inference reveals.
Abstract
AlphaFold2's 93 million parameters, shaped by the evolutionary record of protein structure encoded in the Protein Data Bank and in sequence alignments, are conventionally treated only as machinery for converting sequence to structure. We propose they are also a scientific object that can be analyzed directly: a learned encoding of protein conformational organization that can be probed and characterized. By smoothing the Evoformer's weight tensors with a Gaussian convolution and scaling the result, we show that the trained model produces physically structured conformational landscapes. Under perturbation, ubiquitin's native contacts break in the order established by decades of folding experiments. For KaiB, five independently trained models agree that the alternative fold is not recovered under perturbation. For alpha-synuclein, five models produce five different but coherent landscapes, mapping where the training signal has determined the representation and where it has not. Matched-power noise controls confirm that random corruption of equal magnitude produces debris, not conformations. The model learned to predict static structures; the conformational organization visible under perturbation was not an explicit training target, suggesting it emerged as a byproduct of that objective. AlphaFold2's weights appear to encode structural constraints, shaped by evolutionary and structural training data, that extend beyond what unperturbed inference reveals. We call the approach of reading them neural spectroscopy, and Scaled Gaussian Convolution one such protocol.
Proteins are dynamic molecules capable of adopting multiple conformations. However, AlphaFold2 predominantly generates models around a single conformation, usually representing a ligand-bound state. To address this limitation, we developed AlphaConformers, a structure-guided pipeline that steers AlphaFold2 toward alternative conformations. It is based on the idea that protein structure databases can capture the structural space accessible to members of a protein family. Given a target protein, AlphaConformers retrieves structures from structurally similar proteins. These structures are organized into structure-based alignments and template sets, which are supplied to AlphaFold2 as conformational hypotheses. The resulting models are clustered and filtered, facilitating their analysis. Evaluated on a curated benchmark of 88 proteins with known ligand-bound and unbound conformations, AlphaConformers expanded AlphaFold2 conformational sampling and recovered alternative states missed by AlphaFold2 and other state-of-the-art methods. AlphaConformers ranked first for modelling subtle conformational changes commonly observed between ligand-bound and unbound states. These results show that structural information from protein databases can be leveraged to steer AlphaFold2 toward alternative conformations.
Julie Daniel, Lucas Vitoriano De Queiroz Lira, Diego J. Zea· bioRxiv· 0 citations
Recent advances in de novo protein design have enabled the generation of diverse novel proteins. However, a fundamental challenge remains: even when an amino acid sequence is designed with the target structure as the most stable conformation, there is currently no reliable computational method for assessing whether the target structure is sufficiently stabilized relative to alternative conformations. While experimental realization of the intended fold requires the target structure to be thermodynamically favored by a large free‐energy gap, the absence of a quantitative measure of folding stability makes it difficult to distinguish reliable from unreliable designs. Here, we propose the Residue‐attributed Interpretable Neural network for predicting Absolute folding free energy by Merging structure and sequence Information (RINAMI), a machine learning model that predicts the absolute folding free energy (ΔG) of proteins from their three‐dimensional structures and amino acid sequences. RINAMI integrates structure‐ and sequence‐based representations derived from ProteinMPNN and Evolutionary Scale Modeling 2 (ESM2) using a multi‐head cross‐attention mechanism that contextualizes sequence‐derived signals within the structural environment. Benchmarking RINAMI on both natural and designed proteins from the Mega‐scale and Maxwell datasets shows that it outperforms the tested existing approaches, achieving higher correlations with experimental measurements and improved or comparable prediction errors. An ablation study supports the contribution of sequence–structure integration for predictive accuracy. In addition, RINAMI exhibits strong interpretability by capturing key physicochemical effects, including the destabilizing effect of buried hydrophilic residues, the stabilizing effect of buried hydrophobic residues, and the characteristics of cysteine. Together, these results establish RINAMI as an accurate and interpretable framework for ΔG prediction and provide a practical computational tool for evaluating and prioritizing protein designs prior to experimental testing.
Naoki Tomita, G. Chikenji· Protein Science· 0 citations
The (un)folding rates of natural proteins determine their native stability and functional homeostasis, making them important targets for protein engineering and design. From a prediction standpoint, the rates have been a long‐standing puzzle. We have known for decades that folding rates empirically correlate with properties of the native three dimensional (3D) structures and that both, folding and unfolding rates, scale with protein size. Whereas such rate correlations are too rough for being of practical use, no significant progress in prediction accuracy has occurred since then, despite many efforts even including machine learning approaches. Here, we retake on this challenge by expanding the simple one‐dimensional free energy surface (1D‐FES) model that originally led to demonstrate the size scaling of both rates, and a curated database with rates for 75 single‐domain proteins. We define the weighted sequence order (WSO) as a novel parameter that allows incorporating structural information into the 1D‐FES model explicitly. Via the WSO, we examine the role of global structural properties such as fold topology and core packing in defining the (un)folding rates within the context of a physics‐based model of protein folding. After introducing fold topology and packing at a coarse‐grained level, the model uses three floating parameters to predict the folding and unfolding rates within 6.5‐ and 10‐fold, respectively, resulting in ±6.5 kJ/mol accuracy in native stability, equivalent to the typical perturbation induced by one single‐point mutation. The net improvement over the 2‐parameter size‐only prediction is of 2.5‐fold. These new rate predictions are significantly closer to the threshold of usefulness for engineering and design. More importantly, this WSO‐modified 1D‐FES model can now directly accommodate atomistic, high‐resolution, force‐fields to further optimize the rate predictions, and/or to use rate information as a testbed for force‐field refinement. Finally, the WSO‐1D‐FES model could also serve as foundation for developing more complex models capable of dealing with multi‐domain proteins as well as with the evolutionary information cryptically encoded in natural protein sequences.
Mohammad Abdulqader, Victor Muñoz· Protein Science· 0 citations
Motivation With the rapid development of AI methods that predict protein structures from sequence, understanding the structure-function relation increasingly depends on quantitative structural descriptors that are both biologically meaningful and scalable to large datasets. Here, we introduce mathematical topology metrics that quantify the entanglement complexity of a tertiary protein structure while respecting uncrossability constraints. Results By employing only three such metrics across all protein structures in the Protein Data Bank, we represent the proteome structural space in a continuous three-dimensional space. Distances within this space capture structural similarity and correlate with functional similarity. We find that the mathematical entanglement based landscape of protein structural space diversifies with the evolutionary expansion of protein function across species. Moreover, this continuous representation reproduces CATH classifications with high accuracy for major structural classes. These results indicate that these metrics efficiently encode structural features linked to protein function and provide a more informative description than conventional metrics. Availability Data used in this study are available in the Protein Data Bank. Details of the machine learning model used can be found in https://github.com/roshitac/CATH_Classification-. Contact Banu.Ozkan@asu.edu, Eleni.Panagiotou@asu.edu Supplementary information Supplementary data are available at Journal Name online.
P. Malatesta, Roshita S. Chandnani, J. Yalim et al.· bioRxiv· 0 citations
Solution NMR spectroscopy provides atomistic measurements of proteins in a native-like biophysical state. Because these measurements are ensemble averages, it also has the potential to report on conformational diversity. However, conventional NMR structure determination typically converts experimental observables into restraints for molecular dynamics, which encode information on the mean structure but do not retain information on the underlying conformational distribution. Ensemble selection has long been proposed as an alternative, whereby experimental observables are compared directly with candidate conformers generated independently of the measurements. This allows population distributions to be inferred from the data. However, few such methods have incorporated NOESY - the richest source of structural information in protein NMR - data, due to challenges in the quantitative comparison of experimental and back-calculated spectra. To address this challenge, we previously introduced the CoMAND method, demonstrating that quantitative agreement is practical for NOESY spectra with bespoke heteronuclear editing schemes. Here we extend this approach into a framework for direct inference of protein ensembles within a flexible ensemble-selection architecture incorporating multiple classes of NMR observables. We introduce a quantitative scoring framework for comparing experimental and back-calculated observables and combine it with regularized ensemble selection and Monte Carlo simulated annealing. Integration with the OpenMM molecular dynamics engine allows conformational pools to be generated using established molecular simulation methods. Applied to human ubiquitin, the resulting ensemble provides simultaneous agreement with NOESY, residual dipolar coupling and scalar coupling data while retaining conformational diversity supported by experiment.
A protein’s function is derived from its three-dimensional structure and the motions of the atoms about that structure. The detailed characterization of both macromolecular structure and dynamics provides an opportunity for understanding enzyme catalysis, ligand binding, and allostery, along with providing insights into how the function changes upon mutation or post-translational modification. Among the various methods for characterizing biomolecular motions, nuclear magnetic resonance (NMR) spin relaxation methods are a standard for determining nanosecond global tumbling times along with the amplitude and timescale of faster local motions. Within the model-free formalism, various mathematical models are used to extract dynamic parameters. Unfortunately, as the number of fitted parameters increases within these models, they become mathematically underdetermined for standard NMR relaxation data collected at a single magnetic field, necessitating multi-field datasets. Here, we present SPINDLE, an ensemble of deep neural networks trained on a large synthetic set of NMR relaxation data. Unlike traditional least-squares fitting, SPINDLE predicts both fast and slow timescale dynamics parameters from a single set (i.e., collected at a single magnetic field) of three relaxation datasets using the ensemble for error estimation. We demonstrate a strong correlation to ground truth dynamics parameters on synthetic benchmarks, with more precision than traditional fitting techniques, and precisely reproduce experimental dynamics parameters for ∼50 proteins with relaxation data in the Biological Magnetic Resonance Data Bank. We also leverage the architecture of the deep neural network to show how the model emphasizes rigid residues for the prediction of global correlation times. This strategy may be useful in the future for elucidating correlated networks of dynamic residues from multiple relaxation datasets.
Olivia E. Krise, Michael P. Latham· bioRxiv· 0 citations