Mapping the protein interactome is fundamental to understanding disease mechanisms and facilitating therapeutic development. Although protein language models (PLMs) such as ESM-2 have advanced protein-protein interaction (PPI) prediction, their high-dimensional representations remain difficult to connect to verifiable biological signals. To address this limitation, we propose HybridStack-PPI, a gray-box framework that combines ESM-2 sequence representations with explicit physicochemical and motif-derived biological descriptors. The architecture uses motif-anchored local pooling global mean pooling, symmetric pair encoding, fold-internal feature selection, LightGBM branch learners, and an elastic-net logistic-regression stacking layer. We evaluated the method using a C3 cluster-based cross-validation protocol with a 40% sequence-identity clustering threshold and a Same-GO hard-negative setting in which negative candidates shared functional annotations with positive pairs. Under this setting, HybridStack-PPI reached a Human ROC-AUC of 73.65%, PR-AUC of 91.35%, MCC of 28.06%, and specificity of 75.61%. The results indicate a conservative operating point: compared to more recall-oriented baselines, the proposed stack trades lower recall and F1 for higher specificity, MCC, and ranking behavior under functionally similar negative samples. We further reported cross-species transfer, ablation, latency, SHAP-based descriptor attribution, and meta-learner coefficient analyses to clarify both the promise and limitations of biologically informed PPI prediction.
HyBind-NN is developed, a multimodal graph neural network that integrates protein language models (PLMs) with 3D structural and dynamic datasets to predict protein–protein and protein–peptide affinity, and it is demonstrated that combining ESM-2 sequence embeddings with precise 3D Voronoi spatial geometry enables accurate affinity predictions across diverse structural datasets.
E. A. Bogdanova, A. Chernukhin, Alexey K. Shaytan· International Journal of Mol...· 0 citations
HIPPO (HIerarchical Protein–Protein interaction prediction across Organisms), a hierarchical contrastive learning framework for PPI prediction, is introduced and it is suggested that structured biological knowledge can improve representation learning for PPI prediction across diverse and imbalanced datasets.
Shiyi Liu, Buwen Liang, Yuetong Fang et al.· International Journal of Mol...· 0 citations
The novel combination of LZ complexity–based negative sample selection, CT feature representation, and GA-optimized CNN–LSTM architecture provides a robust and biologically informed framework for PPI prediction.
SPPIPred, an advanced machine learning-based model designed for precise PPI prediction, is presented, offering valuable insights to researchers in the field of bioinformatics and improving applications within bioengineering and pharmaceutical development.
M. Rahman, M. Ali, Md. Shohidullah et al.· PLoS ONE· 0 citations
Accurate prediction of protein-protein interaction interfaces is critical for understanding molecular recognition and guiding therapeutic design. This study presents a comprehensive machine learning pipeline for predicting interface residues in permanent homodimeric protein complexes. Using a curated dataset of 1311 homodimers, we benchmarked six widely used machine learning algorithms and identified multilayer perceptron and XGBoost as top performers, achieving Matthews correlation coefficients (MCC) exceeding 0.93. To enhance interpretability and efficiency, we employed recursive feature elimination to derive a minimal set of six biologically meaningful features, including solvent accessibility, surface roughness, planarity, and average protrusion index, that retained high predictive power (MCC > 0.90). Structurally stratified models tailored to α-helical, β-strand, and membrane proteins demonstrated comparable or improved accuracy relative to generalized models, particularly when utilizing the reduced feature subset. As a preliminary demonstration of generalizability, we applied our approach to an external heterodimer complex (PDB ID: 9ETL). While limited to a single case study, the structurally specialized models maintained high accuracy, suggesting potential applicability beyond the training domain. Furthermore, our residue-level feature-driven models demonstrated highly competitive performance when compared against the baseline established by the general-purpose ColabFold pipeline. The results highlight the importance of structural context in interface prediction and demonstrate that compact, structure-aware models can achieve high accuracy while reducing computational complexity. This work provides a scalable, interpretable, and biologically informed approach to protein interface prediction, with implications for large-scale structural descriptor, drug target characterization, and protein engineering applications.
Tayyip Topuz, Z. Erdem, Halil Bisgin et al.· Scientific Reports· 0 citations
A neural network-based pipeline that integrates amino acid sequences with structural features is developed and provides a modular prototype for follow-up, more extensive protein modeling, including larger proteins and sequence of variable sizes.
Carl David Jasper Causin, M. Fyta· APL Machine Learning· 0 citations