Jul 2026· International Journal of Molecular Sciences· Vol 27, pp. 6242· 0 citations· 43 references
Medicine
TL;DR
HIPPO (HIerarchical Protein–Protein interaction prediction across Organisms), a hierarchical contrastive learning framework for PPI prediction, is introduced and it is suggested that structured biological knowledge can improve representation learning for PPI prediction across diverse and imbalanced datasets.
Abstract
With advances in biomedical technologies and the continued expansion of experimental resources, biological data are growing rapidly in both scale and complexity. Contrastive learning provides an effective framework for integrating heterogeneous biological information. However, many protein–protein interaction (PPI) prediction methods still represent protein sequences and annotations as flat features and do not explicitly model hierarchical biological relationships among protein families, clans, and functional annotations. Here, we introduce HIPPO (HIerarchical Protein–Protein interaction prediction across Organisms), a hierarchical contrastive learning framework for PPI prediction. HIPPO aligns protein sequence representations with structured biological attributes. Across intra-species benchmark PPI datasets, HIPPO improves the average micro-F1 by 2.9% compared with the best baseline across the evaluated splits. In the host–pathogen interaction benchmark, HIPPO achieves the highest AUROC under the standard split (0.731) and the second-best AUPRC (0.332). Under leave-one-virus-family-out evaluation, HIPPO obtains the best AUROC on Papillomaviridae (0.603) and Retroviridae (0.612), while also showing family-dependent transfer behavior. Ablation experiments support the contribution of hierarchical feature integration, and attention-based residue attribution provides preliminary evidence that the learned representations highlight interface-related residues. Together, these results suggest that structured biological knowledge can improve representation learning for PPI prediction across diverse and imbalanced datasets.
Mapping the protein interactome is fundamental to understanding disease mechanisms and facilitating therapeutic development. Although protein language models (PLMs) such as ESM-2 have advanced protein-protein interaction (PPI) prediction, their high-dimensional representations remain difficult to connect to verifiable biological signals. To address this limitation, we propose HybridStack-PPI, a gray-box framework that combines ESM-2 sequence representations with explicit physicochemical and motif-derived biological descriptors. The architecture uses motif-anchored local pooling global mean pooling, symmetric pair encoding, fold-internal feature selection, LightGBM branch learners, and an elastic-net logistic-regression stacking layer. We evaluated the method using a C3 cluster-based cross-validation protocol with a 40% sequence-identity clustering threshold and a Same-GO hard-negative setting in which negative candidates shared functional annotations with positive pairs. Under this setting, HybridStack-PPI reached a Human ROC-AUC of 73.65%, PR-AUC of 91.35%, MCC of 28.06%, and specificity of 75.61%. The results indicate a conservative operating point: compared to more recall-oriented baselines, the proposed stack trades lower recall and F1 for higher specificity, MCC, and ranking behavior under functionally similar negative samples. We further reported cross-species transfer, ablation, latency, SHAP-based descriptor attribution, and meta-learner coefficient analyses to clarify both the promise and limitations of biologically informed PPI prediction.
T. T. Nguyen, X. Mai, N. Nguyen· IEEE Access· 0 citations
Cross-organism prediction of essential proteins is a critical task for drug discovery and microbial engineering, yet the generalizability of existing machine learning models across diverse species remains a significant challenge. In this study, we propose DeepPEP, a large language model-based framework designed to reliably transfer essential protein annotations between distantly related organisms. Utilizing 66 curated prokaryotic datasets, we systematically evaluated DeepPEP's cross-organism performance under various conditions. Initial pairwise predictions revealed a correlation between performance and evolutionary distance; however, further investigation demonstrated that integrating training data from multiple organisms yields superior predictive power. In a benchmark scenario designed to simulate real-world applications, DeepPEP outperformed the state-of-the-art tool Geptop 2.0, showcasing a robust ability to identify species-specific essential proteins. Finally, a case study on novel genomes confirmed the model's practical effectiveness. Our results suggest that DeepPEP is a powerful strategy for prokaryotic essential protein prediction, and the rigorous evaluation framework established in this study provides a new benchmark for the field.
Ying Du, Zhikang Liu, Jing Wan· Journal of Microbiological M...· 0 citations
The utility of HA sites for suggesting candidate binding sites and the biological interpretability of PLM representations is explored, demonstrating the biological interpretability of PLM representations and offers a valuable method to prioritize functionally relevant protein residues for targeted biomedical research.
Sophia J. Pribus, Russ B. Altman, Gowri Nayar· bioRxiv· 0 citations
A neural network-based pipeline that integrates amino acid sequences with structural features is developed and provides a modular prototype for follow-up, more extensive protein modeling, including larger proteins and sequence of variable sizes.
Carl David Jasper Causin, M. Fyta· APL Machine Learning· 0 citations