Skip to content
Open access

X-PAIR: an ultrafast multitask framework for proteome-scale reconstruction of PPI networks and partner-specific interfaces from sequence

Jul 2026 · bioRxiv · 0 citations · 59 references
Biology

TL;DR

X-PAIR is presented, a sequence-based multitask deep learning framework that jointly predicts whether two proteins interact and identifies their partner-specific interface residues, and links proteome-scale interaction discovery to the residue-level determinants of partner-specific molecular recognition.

Abstract

Protein–protein interaction prediction and residue-level interface localisation are biologically intertwined but usually treated as separate computational problems. Here we present X-PAIR, a sequence-based multitask deep learning framework that jointly predicts whether two proteins interact and identifies their partner-specific interface residues. By combining protein language-model representations with lightweight cross-attention, X-PAIR requires neither structural templates nor multiple-sequence alignments. Across leakage-controlled benchmarks, it outperforms existing methods in both tasks, with substantial gains in interface localisation. Multitask learning preserves single-task performance while returning both outputs at near-single-task cost, enabling one million protein pairs to be analysed in under two hours—approximately 500-fold faster for interface prediction and 20-fold faster for PPI prediction than current approaches—thereby enabling proteome-scale analysis. Cross-species analyses reveal distinct evolutionary dependencies: interaction prediction benefits from multispecies training, whereas interface localisation remains robust across taxonomic scales. X-PAIR thus links proteome-scale interaction discovery to the residue-level determinants of partner-specific molecular recognition.

Read PDF

Similar papers

Open access Jul 2026

A Dual- Task Hierarchical Graph Attention Network for Protein-Protein interaction sites Prediction

An innovative two-stage deep learning framework that combines residue-level graph representation learning with protein-level regression to achieve a thorough modeling of protein interactions and gives a better understanding of the structural processes that control PPI.

Oras A. Hussein, E. Al-Shamery · 0 citations
Open access Jul 2026

Expanding the scope of protein language modeling to protein-protein interactions with MSA Pairformer.

Multiple sequence alignment (MSA) Pairformer is presented, a protein language model that builds on AlphaFold2/3's bidirectional refinement between sequence and pairwise residue representations to accurately model the evolution of protein-protein interactions, despite training exclusively on individual chains.

Yo Akiyama, Zhidian Zhang, Olivia Tang et al. · 2 citations
Open access Jul 2026

Dual Cross-Attention Network for Hierarchical Feature Fusion in Protein-Protein Interaction Prediction

Characterizing protein-protein interactions (PPIs) is essential for deciphering core biological processes, including signal transduction, metabolic pathway regulation, immune recognition, and cell cycle control. However, experimental PPI determination remains time-consuming and expensive, driving the adoption of deep learning as an efficient and accurate computational approach. Current deep-learning-based PPI prediction models typically process both intra- and inter-protein as isolated units in feature extraction, thereby ignoring mutual information transfer within a single protein and the interacting pair. To address this limitation, we propose DCAPPI (Dual Cross-Attention network for Protein-Protein Interaction prediction), a novel framework leveraging dual cross-attention modules for hierarchical feature fusion at both intra- and inter-protein levels. First, the Channel Cross-Attention module processes protein sequence and structure as distinct input channels. It generates deep intra-protein representations by performing cross-attention between sequence-derived and structure-derived tokens, achieving multimodal feature integration. Second, the Partner Cross-Attention module models the target protein and its interacting partner as a pair of correlative units. By performing cross-attention operations across these units, it enables collaborative feature fusion and constructs context-aware inter-protein interaction features. Evaluation results indicate that DCAPPI achieves superior performance over state-of-the-art methods on benchmark datasets.

Shuai Lu, Yuguang Li, Zhen Tian et al. · 0 citations
Open access Jul 2026

PCIPG2.0: multi-omics fusion and structure-aware graph autoencoding for protein complex identification

Abstract Motivation Protein complexes execute cellular functions, yet identifying them from protein–protein interaction (PPI) networks remains challenging because interactomes are incomplete and purely topology-driven clustering often lacks mechanistic interpretability. Here we present PCIPG 2.0, an unsupervised framework that explicitly addresses two major bottlenecks in PPI-based complex discovery: missing interactions and limited mechanistic specificity. PCIPG 2.0 first fuses multiple omics views to prioritize high-confidence candidate protein associations and enhance the observed interactome, and then learns structure-aware node representations by aggregating residue embeddings on residue graphs, coupled with a PPI-level graph autoencoder to infer latent complex memberships. Results Across five yeast benchmarks, PCIPG 2.0 consistently improves complex recovery over representative baselines and yields predicted complexes with significantly higher Gene Ontology semantic coherence than size-matched random sets. Literature-supported case studies and AlphaFold3-based assembly analyses further suggest that representative predictions are consistent with functionally coherent and structurally plausible protein assemblies. Together, these results suggest that combining multi-omics-driven interactome completion with residue-informed representation learning provides a useful and mechanistically informed framework for protein complex identification under incomplete interactome measurements. Availability The source code is available at GitHub: https://github.com/hyx-1/PCIPG2.0. The complete reproducibility package, including the code, processed data, configuration files and materials required to reproduce the experiments reported in this manuscript, has been archived on Zenodo with an archival DOI: https://doi.org/10.5281/zenodo.20133228. The GitHub repository also provides the README-based usage instructions.

Yixiang Huang, Jiudong Wang, Lei Yang et al. · 0 citations
Open access Jul 2026

Predictions of protein–protein interactions: Learning sequences and structures

A neural network-based pipeline that integrates amino acid sequences with structural features is developed and provides a modular prototype for follow-up, more extensive protein modeling, including larger proteins and sequence of variable sizes.

Carl David Jasper Causin, M. Fyta · 0 citations
Open access Aug 2026

PLMView: collaborative protein language model representations for fast and scalable specialized protein function inference

Applications to thioredoxins, visual opsins, and Tara Oceans environmental diatom cold-shock proteins show that PLMView can move from interpretable residue-level determinants in well-studied protein families to large-scale environmental functional discovery, linking molecular specialization to ecological distribution and transcriptional deployment across the global ocean.

Vinh-Son Pho, Alessandro Natale Bianchi, Mattéo Scarsini et al. · 0 citations