X-PAIR is presented, a sequence-based multitask deep learning framework that jointly predicts whether two proteins interact and identifies their partner-specific interface residues, and links proteome-scale interaction discovery to the residue-level determinants of partner-specific molecular recognition.
Abstract
Protein–protein interaction prediction and residue-level interface localisation are biologically intertwined but usually treated as separate computational problems. Here we present X-PAIR, a sequence-based multitask deep learning framework that jointly predicts whether two proteins interact and identifies their partner-specific interface residues. By combining protein language-model representations with lightweight cross-attention, X-PAIR requires neither structural templates nor multiple-sequence alignments. Across leakage-controlled benchmarks, it outperforms existing methods in both tasks, with substantial gains in interface localisation. Multitask learning preserves single-task performance while returning both outputs at near-single-task cost, enabling one million protein pairs to be analysed in under two hours—approximately 500-fold faster for interface prediction and 20-fold faster for PPI prediction than current approaches—thereby enabling proteome-scale analysis. Cross-species analyses reveal distinct evolutionary dependencies: interaction prediction benefits from multispecies training, whereas interface localisation remains robust across taxonomic scales. X-PAIR thus links proteome-scale interaction discovery to the residue-level determinants of partner-specific molecular recognition.
An innovative two-stage deep learning framework that combines residue-level graph representation learning with protein-level regression to achieve a thorough modeling of protein interactions and gives a better understanding of the structural processes that control PPI.
Oras A. Hussein, E. Al-Shamery· Journal of Intelligent Infor...· 0 citations
Multiple sequence alignment (MSA) Pairformer is presented, a protein language model that builds on AlphaFold2/3's bidirectional refinement between sequence and pairwise residue representations to accurately model the evolution of protein-protein interactions, despite training exclusively on individual chains.
Yo Akiyama, Zhidian Zhang, Olivia Tang et al.· Cell· 2 citations
Characterizing protein-protein interactions (PPIs) is essential for deciphering core biological processes, including signal transduction, metabolic pathway regulation, immune recognition, and cell cycle control. However, experimental PPI determination remains time-consuming and expensive, driving the adoption of deep learning as an efficient and accurate computational approach. Current deep-learning-based PPI prediction models typically process both intra- and inter-protein as isolated units in feature extraction, thereby ignoring mutual information transfer within a single protein and the interacting pair. To address this limitation, we propose DCAPPI (Dual Cross-Attention network for Protein-Protein Interaction prediction), a novel framework leveraging dual cross-attention modules for hierarchical feature fusion at both intra- and inter-protein levels. First, the Channel Cross-Attention module processes protein sequence and structure as distinct input channels. It generates deep intra-protein representations by performing cross-attention between sequence-derived and structure-derived tokens, achieving multimodal feature integration. Second, the Partner Cross-Attention module models the target protein and its interacting partner as a pair of correlative units. By performing cross-attention operations across these units, it enables collaborative feature fusion and constructs context-aware inter-protein interaction features. Evaluation results indicate that DCAPPI achieves superior performance over state-of-the-art methods on benchmark datasets.
Shuai Lu, Yuguang Li, Zhen Tian et al.· Computational and Structural...· 0 citations
Abstract Motivation Protein complexes execute cellular functions, yet identifying them from protein–protein interaction (PPI) networks remains challenging because interactomes are incomplete and purely topology-driven clustering often lacks mechanistic interpretability. Here we present PCIPG 2.0, an unsupervised framework that explicitly addresses two major bottlenecks in PPI-based complex discovery: missing interactions and limited mechanistic specificity. PCIPG 2.0 first fuses multiple omics views to prioritize high-confidence candidate protein associations and enhance the observed interactome, and then learns structure-aware node representations by aggregating residue embeddings on residue graphs, coupled with a PPI-level graph autoencoder to infer latent complex memberships. Results Across five yeast benchmarks, PCIPG 2.0 consistently improves complex recovery over representative baselines and yields predicted complexes with significantly higher Gene Ontology semantic coherence than size-matched random sets. Literature-supported case studies and AlphaFold3-based assembly analyses further suggest that representative predictions are consistent with functionally coherent and structurally plausible protein assemblies. Together, these results suggest that combining multi-omics-driven interactome completion with residue-informed representation learning provides a useful and mechanistically informed framework for protein complex identification under incomplete interactome measurements. Availability The source code is available at GitHub: https://github.com/hyx-1/PCIPG2.0. The complete reproducibility package, including the code, processed data, configuration files and materials required to reproduce the experiments reported in this manuscript, has been archived on Zenodo with an archival DOI: https://doi.org/10.5281/zenodo.20133228. The GitHub repository also provides the README-based usage instructions.
Yixiang Huang, Jiudong Wang, Lei Yang et al.· Bioinformatics· 0 citations
A neural network-based pipeline that integrates amino acid sequences with structural features is developed and provides a modular prototype for follow-up, more extensive protein modeling, including larger proteins and sequence of variable sizes.
Carl David Jasper Causin, M. Fyta· APL Machine Learning· 0 citations
Applications to thioredoxins, visual opsins, and Tara Oceans environmental diatom cold-shock proteins show that PLMView can move from interpretable residue-level determinants in well-studied protein families to large-scale environmental functional discovery, linking molecular specialization to ecological distribution and transcriptional deployment across the global ocean.