Skip to content
Open access

A map of human protein-protein interaction embeddings for functional discovery

Aug 2026 · bioRxiv · 0 citations · 45 references
Biology

Abstract

A protein’s function depends not just on its own structure and localization, but also on the interactions with its partners. Many proteins are therefore better described by a set of partner-dependent roles than by a single annotation. Yet most approaches to the functional interpretation of protein-protein interactions (PPIs) remain protein or set-centric. They rely on pre-existing annotations, and perform worst where knowledge is sparse. Here, we present MAPPIE (Map of Protein-Protein Interaction Embeddings), a method that treats each PPI, rather than each protein, as a unit of representation. From 199,137 human interactions spanning 15,503 proteins, we build a two-dimensional map of the human PPI landscape for functional discovery. Protein language model embeddings for two protein interaction partners are combined and compressed into a latent space, with model selection guided by domain-domain interactions used as a structural proxy for interaction similarity. The resulting geometry separates domain defined interaction classes, organizes disorder associated interactions spatially, and splits interactions involving the same protein by partner. A query PPI’s latent neighbourhood recovers its own annotated functions across molecular, complex, pathway, and biological processes. MAPPIE contributes most where existing functional evidence is weakest, outperforming interactome and sequence identity baselines for sparsely connected interactions. MAPPIE neighbours of query PPIs are enriched for partners in independent protein networks, recovering curated complex-level function even when subunits are spread across the map. Applied to a human dark interactome, MAPPIE assigns specific, experimentally supported functions to dark hub proteins.

Read PDF

Similar papers

Open access Jul 2026

Predictions of protein–protein interactions: Learning sequences and structures

A neural network-based pipeline that integrates amino acid sequences with structural features is developed and provides a modular prototype for follow-up, more extensive protein modeling, including larger proteins and sequence of variable sizes.

Carl David Jasper Causin, M. Fyta · 0 citations
Open access Jul 2026

A Dual- Task Hierarchical Graph Attention Network for Protein-Protein interaction sites Prediction

An innovative two-stage deep learning framework that combines residue-level graph representation learning with protein-level regression to achieve a thorough modeling of protein interactions and gives a better understanding of the structural processes that control PPI.

Oras A. Hussein, E. Al-Shamery · 0 citations
Open access Aug 2026

PLMView: collaborative protein language model representations for fast and scalable specialized protein function inference

Applications to thioredoxins, visual opsins, and Tara Oceans environmental diatom cold-shock proteins show that PLMView can move from interpretable residue-level determinants in well-studied protein families to large-scale environmental functional discovery, linking molecular specialization to ecological distribution and transcriptional deployment across the global ocean.

Vinh-Son Pho, Alessandro Natale Bianchi, Mattéo Scarsini et al. · 0 citations
Open access Jul 2026

Expanding the scope of protein language modeling to protein-protein interactions with MSA Pairformer.

Multiple sequence alignment (MSA) Pairformer is presented, a protein language model that builds on AlphaFold2/3's bidirectional refinement between sequence and pairwise residue representations to accurately model the evolution of protein-protein interactions, despite training exclusively on individual chains.

Yo Akiyama, Zhidian Zhang, Olivia Tang et al. · 2 citations
Open access 2026

A Biologically Informed Hybrid Stacking Framework for Protein–Protein Interaction Prediction

Mapping the protein interactome is fundamental to understanding disease mechanisms and facilitating therapeutic development. Although protein language models (PLMs) such as ESM-2 have advanced protein-protein interaction (PPI) prediction, their high-dimensional representations remain difficult to connect to verifiable biological signals. To address this limitation, we propose HybridStack-PPI, a gray-box framework that combines ESM-2 sequence representations with explicit physicochemical and motif-derived biological descriptors. The architecture uses motif-anchored local pooling global mean pooling, symmetric pair encoding, fold-internal feature selection, LightGBM branch learners, and an elastic-net logistic-regression stacking layer. We evaluated the method using a C3 cluster-based cross-validation protocol with a 40% sequence-identity clustering threshold and a Same-GO hard-negative setting in which negative candidates shared functional annotations with positive pairs. Under this setting, HybridStack-PPI reached a Human ROC-AUC of 73.65%, PR-AUC of 91.35%, MCC of 28.06%, and specificity of 75.61%. The results indicate a conservative operating point: compared to more recall-oriented baselines, the proposed stack trades lower recall and F1 for higher specificity, MCC, and ranking behavior under functionally similar negative samples. We further reported cross-species transfer, ablation, latency, SHAP-based descriptor attribution, and meta-learner coefficient analyses to clarify both the promise and limitations of biologically informed PPI prediction.

T. T. Nguyen, X. Mai, N. Nguyen · 0 citations