Skip to content

Unveiling Large-Scale Kinase-Centric Protein-Protein Interactions through a Knowledge-Informed Workflow.

Jul 2026 · Journal of Chemical Information and Modeling · 0 citations · 44 references
Medicine

TL;DR

A pipeline reformulating kinase-substrate modeling as a Bayesian inference problem is presented and it is revealed that the interaction types and distances to the catalytic pocket significantly influence pathogenicity scores.

Abstract

Protein phosphorylation regulates signaling, yet atomic-level substrate specificity remains elusive due to sparse structural data and phosphorylation-site-insensitive deep-learning predictors. Here we present a pipeline reformulating kinase-substrate modeling as a Bayesian inference problem. By integrating curated data sets and literature evidence parsed by Large Language Models, we converted diverse biological knowledge into structural restraints for the restraint-guided deep-learning model GRASP. For EGFR, BRAF and JNK1, we obtained 336 new phosphorylation-site-specific structure candidates refined by molecular dynamics. These models recapitulate known features, such as JNK1's hydrophobic docking groove, and enabled a Virtual Position Scanning Peptide Array (V-PSPA) to map recognition patches and derive sequence preferences. Cross-referencing predicted interfaces with AlphaMissense pathogenicity scores reveal that the interaction types and distances to the catalytic pocket significantly influence pathogenicity scores. A comparison with clinical mutation data sets further connects pathogenic mutations to the kinase-substrate interface. This high-resolution, high-throughput pipeline can be broadly applicable to kinase specificity studies and general drug discovery.

View source

Similar papers

Open access Aug 2026

Structure-conditioned self-supervised learning of residue interaction constraints in protein kinases for variant interpretation

Abstract Motivation Protein kinases are key regulators of cellular signaling and are frequently implicated in human diseases. Although kinase domains are structurally conserved, predicting the effects of amino acid substitutions remains challenging as mutations often introduce subtle structural perturbations that are not captured by sequence-based or evolutionary methods. Existing supervised approaches further rely on pathogenicity annotations that are inconsistent across databases, thereby motivating the development of structure-based, label-independent frameworks for mutation effect prediction. Results We present a structure-based method using SE(3)-transformers to learn residue compatibility with the local structural environment from experimentally resolved kinase 3D structures. Proteins are represented as atom-level graphs with physicochemical descriptors derived from the CHARMM force field and spatial connectivity. The model is trained on two self-supervised tasks given local structural context: masked residue atom reconstruction and masked residue classification. This formulation enables learning of geometric and physicochemical constraints without relying on pathogenicity labels. Evaluation using reconstruction loss, residue prediction accuracy, and comparison with BLOSUM substitution patterns indicate that the model captures biologically meaningful relationships between residue identity and 3D structural context. We interpret the scores assigned to alternative amino acids as measures of structural fitness, where low-scoring residues are hypothesized to be less compatible with the local environment and more likely to induce deleterious effects on protein structure and activity. Availability and implementation https://zenodo.org/records/20393799.

Shakiba Fadaei, F. Krebs, V. Zoete · 0 citations
Book Open access Aug 2026

CAHS: Context-Aware Homology Search

Protein homology search is foundational to bioinformatics: it supports annotation transfer, structure/function inference, and evolutionary analysis over rapidly expanding sequence repositories (e.g., UniProtKB). Profile hidden Markov models (pHMMs), as implemented in HMMER, remain the most widely trusted approach because they provide statistically calibrated E-values; however, their gap behavior is fixed once a profile is trained, despite biological evidence that insertion/deletion tolerance varies across flexible loops and intrinsically disordered regions. We present CAHS (Context-Aware Homology Search), a lightweight query-time adapter for pHMM search that incorporates learned and biologically motivated signals without changing HMMER's downstream search pipeline or its calibrated E-value reporting. Given a query sequence, CAHS computes per-residue representations from a protein language model and a disorder predictor, maps these to profile coordinates, and modulates only match-state transition rows (gap-open and gap-extension probabilities) while preserving Plan7 constraints. We comprehensively evaluate CAHS across six structurally diverse protein families and multi-domain architectures against a 570k-sequence target corpus. CAHS expands detection capability, retrieving thousands of additional remote homologs at relaxed thresholds by maintaining alignment quality through flexible regions. For multi-domain proteins, context-aware modulation resolves 94% of fragmented alignments. Crucially, CAHS preserves hit-set invariance at stringent operating points (E<10-10), demonstrating increased statistical confidence without inflating false positives. Furthermore, sharper statistical distinction between homologs and background noise during early filter stages yields up to a 3.87× acceleration in end-to-end wall-clock time on high-performance computing clusters. Overall, CAHS illustrates a practical AI-for-science design pattern: augmenting a trusted probabilistic model with query-specific learned signals to improve interpretable, reproducible inference in data-rich biology.

Swethasree Bhattaram, D. Bhowmik, Ramakrishnan Kannan · 0 citations
Review 2026

AI-Driven Protein Research: From Prediction to Design.

This mini review traces the evolution of AI-driven methods in protein research, from early residue-contact prediction using coevolutionary information to transformative breakthroughs, the rise of protein language models (PLMs), and the emerging era of generative design and functional modeling.

Guodong Min, Huan Peng · 0 citations
Open access Jul 2026

Minimal Data · Maximal Insight (MDMI): A Structure-guided Pipeline for Discovering Functional Alternatives in Peptide-Protein Interfaces

Minimal Data Maximal Insight (MDMI), a two-stage structure-guided computational pipeline that designs functional peptide variants using only a small, annotated dataset, demonstrates that structure-informed pipelines can uncover remote functional sequence space from minimal data.

P. Bayat, Spencer J. Perkins, Sebastian Clancy et al. · 0 citations
Open access Jul 2026

Scop3P in 2026: an expanded proteomics-informed resource contextualizing phosphorylation sites through sequence, structure, mutation, and experimental provenance

Protein phosphorylation is a central regulatory mechanism controlling protein activity, interactions, and cellular signalling, and its dysregulation is implicated in numerous diseases. Advances in mass spectrometry–based phosphoproteomics have led to a rapid expansion in the number of reported phosphorylation sites; however, interpretation of these data remains challenging due to fragmented evidence, limited structural context, and the lack of uniform experimental provenance across resources. Interpretation is further complicated by the fact that the biological meaning of reported phosphosites can vary substantially across tissues, cell lines, perturbations, and disease settings. Here, we present a major update of Scop3P, a proteomics-informed knowledgebase that contextualizes human phosphorylation sites within integrated sequence, structural, biophysical, evolutionary, and mutational frameworks. The current release incorporates uniformly reprocessed human phosphoproteomics data from 116 PRIDE datasets alongside curated UniProt annotations, retaining peptide-spectrum matches, site localization confidence, and direct links to primary mass spectrometry evidence via Universal Spectrum Identifiers. This integration yields 152,350 unique serine, threonine, and tyrosine phosphorylation sites across 16,533 human proteins, supported by full experimental provenance. Beyond site identification, Scop3P provides residue-level contextual annotations derived from experimentally determined protein structures and proteome-wide AlphaFold models, enabling near-complete structural coverage of phosphorylation sites. Structural context is further complemented by residue-level biophysical, evolutionary, and mutational annotations, supporting integrated assessment of phosphorylation in functional and disease-related settings. The current release also introduces residue interaction network representations derived from AlphaFold-predicted structures, capturing spatial connectivity and local interaction environments of phosphorylation and mutation sites. A redesigned web interface enables interactive exploration through coordinated 1D, 2D, 2.5D, and 3D visualizations, peptide-level coverage views, and direct access to original spectra via PRIDE. By bridging experimental phosphoproteomics with structural, functional, and disease-related context, Scop3P provides a scalable and provenance-aware resource for phosphosite interpretation, hypothesis generation, and data-driven modelling of phosphorylation-dependent regulation.

P. Ramasamy, Natalia Tichshenko, Adrián Díaz et al. · 0 citations