Skip to content
Open access

PCIPG: A comprehensive framework for protein complex identification based on a probabilistic graphical model

Sep 2026 · PLoS Computational Biology · Vol 22, pp. e1014698 · 0 citations · 83 references
Medicine

TL;DR

PCIPG is presented, a multi-scale probabilistic graph framework that jointly models residues, proteins, interactions and complexes and bridges residue-scale structural cues with interactome-scale organization to enable interpretable and scalable protein complex identification.

Abstract

Protein complexes are molecular machines that execute essential cellular functions, but their computational identification remains challenging. Existing protein complex identification methods largely rely on PPI network topology, functional annotations, or protein-level biochemical evidence. Although these approaches have recovered many biologically meaningful assemblies, they are often sensitive to incomplete or noisy interactomes and provide limited mechanistic insight into the residue- and interface-level determinants of complex formation. In particular, conventional PPI-based graph representations indicate whether proteins are associated, but usually ignore how protein subunits physically interact through spatially organized residues and structural interfaces. These limitations motivate the development of computational frameworks that connect residue-scale structural cues with interactome-scale organization. Here we present PCIPG, a multi-scale probabilistic graph framework that jointly models residues, proteins, interactions and complexes. PCIPG encodes residue-level physicochemical descriptors on intra-chain contact maps, screens informative residues to construct structure-aware protein representations and propagates these representations over the PPI graph to infer a protein–complex membership matrix. To couple complex membership with sparse interaction evidence, PCIPG reconstructs the network using a zero-inflated Bernoulli–Exponential likelihood, providing a principled learning signal under missing-edge and noise regimes. Across five Saccharomyces cerevisiae benchmarks, PCIPG achieved higher average F1 and Acc than the representative baseline methods included in this study, with average improvements of 11.46% and 3.64%, respectively. On the evaluated human interactomes, PCIPG achieved the highest F1 score among the compared methods on HCT116 and HEK293T, whereas its performance on HuRI was below that of AdaPPI and ClusterONE. Embedding-guided interaction completion improved PCIPG’s performance relative to its results on the corresponding original human PPI networks. Beyond complex calling, PCIPG supports core–module mining by recovering known cores and delineating coherent accessory modules within assemblies; several predictions match previously reported functional entities, including TRAPPII- and PCNA-loading-factor–related complexes. At the residue level, residues prioritized by PCIPG show increased overlap with experimentally defined protein-binding interfaces in the evaluated structures. In a computational CFTR case study, the model generated state-dependent interaction predictions that partially overlapped with experimentally profiled wild-type and ΔF508 interaction networks. Together, PCIPG bridges residue-scale structural cues with interactome-scale organization to enable interpretable and scalable protein complex identification. Code and data are available at https://github.com/hyx-1/PCIPG.

Read PDF

Similar papers

#protein folding Open access Sep 2026

LoGoPPI enables fast and accurate protein–protein interaction mapping at scale

Graph-based protein function analysis is powerful, but protein-protein interaction (PPI) networks exist for only a small fraction of animal and plant genomes. We present LoGoPPI, which infers PPIs from sequence by combining bi-encoder global protein representation with local residue-level late interaction. LoGoPPI matc...

Hae Been Lee, Junyeong Ma, Han-June Kim et al. · 0 citations
Open access Sep 2026

Site-resolved spatial and structural interactome of a human cell

The spatial and structural arrangement of proteins determine virtually every process in human cells. We combined gentle subcellular fractionation by differential ultracentrifugation with cross-linking mass spectrometry to systematically map this cellular proteome architecture with residue-level evidence, identifying 16...

Ze-Hong Zhang, Nanako Yokoyama, Ying Zhu et al. · 0 citations
Open access Sep 2026

Multi-vector retrieval enables residue-resolved prediction of protein partners by ColBERT-PPI

Protein–protein interactions (PPIs) are central to biological processes, making the identification of both interacting partners and their binding sites important for understanding molecular function and guiding therapeutic discovery. However, connecting large-scale partner prediction to residue-level interaction eviden...

He Yang, Ru-Xin Lei, You-Wen Zhuang et al. · 0 citations
Review Sep 2026

Proximity labeling: A powerful tool for mapping protein interactions.

Protein-protein interactions (PPIs) are fundamental to cellular regulation, and their dysregulation contributes to numerous diseases. Conventional approaches for studying PPIs often lack sufficient spatial and temporal resolution and are limited in capturing weak, transient, or context-dependent molecular associations....

Hua Yin, Shi-Wu Zhang, Shu-Lin Liu et al. · 0 citations
Open access Sep 2026

The human metabolite–protein interactome reveals a global layer of cellular coordination

Metabolites are substrates, products, cofactors, and regulators, but protein–protein interaction networks do not represent their potential to organize proteins across conventional pathway boundaries. Using the LIGMAP virtual-screening algorithm, we mapped 308 human metabolite codes to pockets in monomers, dimer interfa...

J. Skolnick, B. Srinivasan · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.