PCIPG is presented, a multi-scale probabilistic graph framework that jointly models residues, proteins, interactions and complexes and bridges residue-scale structural cues with interactome-scale organization to enable interpretable and scalable protein complex identification.
Abstract
Protein complexes are molecular machines that execute essential cellular functions, but their computational identification remains challenging. Existing protein complex identification methods largely rely on PPI network topology, functional annotations, or protein-level biochemical evidence. Although these approaches have recovered many biologically meaningful assemblies, they are often sensitive to incomplete or noisy interactomes and provide limited mechanistic insight into the residue- and interface-level determinants of complex formation. In particular, conventional PPI-based graph representations indicate whether proteins are associated, but usually ignore how protein subunits physically interact through spatially organized residues and structural interfaces. These limitations motivate the development of computational frameworks that connect residue-scale structural cues with interactome-scale organization. Here we present PCIPG, a multi-scale probabilistic graph framework that jointly models residues, proteins, interactions and complexes. PCIPG encodes residue-level physicochemical descriptors on intra-chain contact maps, screens informative residues to construct structure-aware protein representations and propagates these representations over the PPI graph to infer a protein–complex membership matrix. To couple complex membership with sparse interaction evidence, PCIPG reconstructs the network using a zero-inflated Bernoulli–Exponential likelihood, providing a principled learning signal under missing-edge and noise regimes. Across five Saccharomyces cerevisiae benchmarks, PCIPG achieved higher average F1 and Acc than the representative baseline methods included in this study, with average improvements of 11.46% and 3.64%, respectively. On the evaluated human interactomes, PCIPG achieved the highest F1 score among the compared methods on HCT116 and HEK293T, whereas its performance on HuRI was below that of AdaPPI and ClusterONE. Embedding-guided interaction completion improved PCIPG’s performance relative to its results on the corresponding original human PPI networks. Beyond complex calling, PCIPG supports core–module mining by recovering known cores and delineating coherent accessory modules within assemblies; several predictions match previously reported functional entities, including TRAPPII- and PCNA-loading-factor–related complexes. At the residue level, residues prioritized by PCIPG show increased overlap with experimentally defined protein-binding interfaces in the evaluated structures. In a computational CFTR case study, the model generated state-dependent interaction predictions that partially overlapped with experimentally profiled wild-type and ΔF508 interaction networks. Together, PCIPG bridges residue-scale structural cues with interactome-scale organization to enable interpretable and scalable protein complex identification. Code and data are available at https://github.com/hyx-1/PCIPG.
Graph-based protein function analysis is powerful, but protein-protein interaction (PPI) networks exist for only a small fraction of animal and plant genomes. We present LoGoPPI, which infers PPIs from sequence by combining bi-encoder global protein representation with local residue-level late interaction. LoGoPPI matc...
Hae Been Lee, Junyeong Ma, Han-June Kim et al.· bioRxiv· 0 citations
The spatial and structural arrangement of proteins determine virtually every process in human cells. We combined gentle subcellular fractionation by differential ultracentrifugation with cross-linking mass spectrometry to systematically map this cellular proteome architecture with residue-level evidence, identifying 16...
Protein–protein interactions (PPIs) are central to biological processes, making the identification of both interacting partners and their binding sites important for understanding molecular function and guiding therapeutic discovery. However, connecting large-scale partner prediction to residue-level interaction eviden...
He Yang, Ru-Xin Lei, You-Wen Zhuang et al.· bioRxiv· 0 citations
Protein-protein interactions (PPIs) are fundamental to cellular regulation, and their dysregulation contributes to numerous diseases. Conventional approaches for studying PPIs often lack sufficient spatial and temporal resolution and are limited in capturing weak, transient, or context-dependent molecular associations....
Hua Yin, Shi-Wu Zhang, Shu-Lin Liu et al.· Biosensors & bioelectronics· 0 citations
Metabolites are substrates, products, cofactors, and regulators, but protein–protein interaction networks do not represent their potential to organize proteins across conventional pathway boundaries. Using the LIGMAP virtual-screening algorithm, we mapped 308 human metabolite codes to pockets in monomers, dimer interfa...
J. Skolnick, B. Srinivasan· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.