Decoding the allosteric grammar of protein kinases: A dual‐stream framework integrating protein language models and energy landscape frustration analysis
Jul 2026· Protein Science· Vol 35· 1 citation· 98 references
Medicine
TL;DR
This study reveals how the organization of the protein energy landscape shapes universal "allosteric grammar" and algorithmic detectability of regulatory binding sites and proposes that allosteric sites are encoded in persistent neutrally frustrated regions optimized for context‐dependent regulatory modulation.
Abstract
The spatial and energetic encoding of allosteric regulatory sites remains a major challenge in structural biology, frequently representing a “blind spot” for sequence‐based artificial intelligence (AI) models. We present a protein language model (PLM)‐guided approach complemented by the energy landscape frustration analysis as a dual‐stream framework to investigate the relationship between AI prediction of binding sites and biophysical organization of regulatory pockets across the human kinome. By probing a fine‐tuned residue‐level PLM classifier across 453 kinase structures, a clear performance gap is discovered between highly predictable orthosteric pockets (Types I, I.5, and II) and poorly resolved distal allosteric sites (Type IV). Rather than attempting to interpret this blind spot through internal AI attributions alone, we use independent local frustration profiles to analyze the underlying physics of these sites. We determine that the detectability of orthosteric and allosteric binding sites reflects their energetic embedding within the protein energy landscape. Orthosteric catalytic sites reside within minimally frustrated, optimized energetic regions that are consistently detected with high confidence. In contrast, allosteric sites are enriched in neutrally frustrated zones, producing diffuse and context‐dependent predictions. We demonstrate that this neutral frustration of functional regions acts as a biophysical lubricant, facilitating the conformational plasticity required for regulatory transitions while simultaneously eroding the coevolutionary signals exploited by PLMs. Atomic‐resolution analysis of abelson murine leukemia (ABL) kinase spanning multiple conformational states and complexes bound to diverse ligands provides mechanistic validation of this principle. The myristoyl allosteric pocket in ABL remains neutrally frustrated across complexes with physiological ligands, chemically diverse modulators, from allosteric inhibitors to activators, and conformations engaged with SH2–SH3 regulatory domains. We propose that allosteric sites are encoded in persistent neutrally frustrated regions optimized for context‐dependent regulatory modulation. This study reveals how the organization of the protein energy landscape shapes universal “allosteric grammar” and algorithmic detectability of regulatory binding sites.
Allosteric modulation of G protein–coupled receptors (GPCRs) offers major advantages in receptor selectivity and signaling control; yet systematic approaches to identify allosteric modulators, define their binding sites, and map the underlying allosteric networks remain limited. Current molecular dynamics (MD) and machine learning (ML)-based methods often rely on correlation-driven or black-box models that provide limited mechanistic insight. We developed an interpretable probabilistic framework that extracts residue-level dependencies from MD ensembles using Bayesian network modeling (BNM). By representing each residue through its local interaction energy, BNM identifies both local and long-range energetic couplings and maps the allosteric communication pathways linking the AngII binding site to the G-protein interface in the angiotensin II type 1 receptor (AT1R). To functionally prioritize these pathways, we integrated BNM with comprehensive mutational analysis, combining whole-receptor alanine mutagenesis data with exhaustive in silico deep mutational scanning to validate BNM-predicted hotspots. This approach recovered state-dependent allosteric communities, revealed residues in noncanonical regions that regulate Gαq coupling and identified positions whose functional importance emerged only with specific, predicted substitutions, as well as highlighted a cryptic intracellular pocket enriched in communication hubs. Guided by these network-derived residues and pocket geometries, structure-based virtual screening identified a small, fragment-like molecule negative allosteric modulator (NAM) named Q2 that attenuates AngII-mediated Gαq signaling. Mutational mapping supports Q2 binding adjacent to the G-protein interface, consistent with its mechanism of action. Together, these results establish a generalizable and interpretable framework for uncovering GPCR allosteric communication networks and discovering modulators that exploit these networks.
It is argued that incorporating frustration into computational and experimental strategies will be essential to move beyond purely stability-driven approaches toward the rational engineering of functional proteins.
Franco L. Simonetti, Eli J. Draizen, Rocío Espada et al.· Biochimica et Biophysica Act...· 0 citations
Allostery is increasingly understood as the propagation of dynamic information through a protein, yet the computational descriptors of that flow are computed one structure at a time and carry no transferable, sequence-level prior. Here we build an alphabet of allostery: a dictionary of local contact words, short sequence windows anchored by three-to four-residue spatial cliques, each carrying a distribution of net Gaussian network model transfer entropy scores pooled over a non-redundant set of Protein Data Bank structures. The alphabet consists of 131,611,766 unique words drawn from 212,860,934 clique observations. Projecting any protein’s sequence and structure onto this dictionary yields a per-residue allosteric track, net TE, source, sink, and switch channels, with no system-specific fitting. Validating against the Allosteric Database, annotated allosteric-site residues behave as transfer entropy sinks (information receivers): the sink channel discriminates sites from the rest of the protein with pooled ROC-AUC = 0.543 over 646,629 residues (permutation z = 19.1), an effect small in magnitude but overwhelmingly significant and robust to word-frequency leakage. The directional channels are mechanistically informative in a two-state experiment on nine canonical allosteric proteins: source residues predict the largest apo→holo conformational rewiring (meta-analytic Spearman ρ = +0.106, positive in 7/9 proteins) while sink residues mark the most conformationally stable positions (ρ = −0.105, 8/9). Sinks thus mark where allosteric signal is received and sources mark where it drives motion. Finally, we mine the most context-variable words into a compact, hydrophobic-enriched switch vocabulary that we propose as a design dictionary for engineering allosteric mechanisms.
We survey modern deep‐learning approaches to protein conformational modeling through the lens of architectural design. We organize the literature into three increasingly expressive paradigms: (I) single‐structure prediction, (II) prediction of molecular binding complexes, and (III) conformational ensemble generation. For each paradigm, we outline a representative set of models to sketch a practical taxonomy, and we summarize their key achievements, limitations, and common evaluation practices. Across the paradigms, we highlight recurring design choices that shape performance and generalization, including enforced SE(3) equivariance versus learned symmetry; MSA‐driven coevolution versus protein language model priors; deterministic prediction versus generative sampling; explicit energetic supervision versus implicit learning; and integrative modeling across heterogeneous data modalities. While single‐structure prediction is now relatively well established, comparable maturity has not yet been reached for binding‐complex prediction and, especially, for generating faithful thermodynamic ensembles with reliable population weights, which remains an open challenge. We discuss open challenges in building physically grounded and transferable models, including data availability and fidelity, the choice of inductive biases to pursue generalization, and the need for rigorous model evaluation. Ultimately, we indicate generative kinetics as an aspirational frontier.
Daniele Angioletti, Matteo Carli, Marco S. Nobile et al.· WIREs Computational Molecula...· 0 citations
Results show that current co-folding methods remain unreliable as stand-alone predictors of ion-channel ligand-binding modes and highlight pose sampling, pocket selection, ligand representation and independent structural validation as priorities for methodological development.
Yu Zhu, Taufiq Rahman· Frontiers in Biophysics· 0 citations
PHASE (Protein Hamiltonians for Sampling of Ensembles), a system-specific framework that converts atomistic conformational ensembles into an explicit and interpretable statistical model, is introduced.
Daniele Angioletti, Marco S. Nobile, Matteo Carli et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.