Skip to content
Open access

SPPIPred: Stacking-based ensemble learning model for identification of protein-protein interaction

Jul 2026 · PLoS ONE · Vol 21 · 0 citations · 72 references
Medicine

TL;DR

SPPIPred, an advanced machine learning-based model designed for precise PPI prediction, is presented, offering valuable insights to researchers in the field of bioinformatics and improving applications within bioengineering and pharmaceutical development.

Abstract

Protein-protein interactions (PPIs) are essential for various biological functions and are crucial in drug discovery, signaling pathways, and network reconstruction. This study presents SPPIPred, an advanced machine learning-based model designed for precise PPI prediction. The SPPIPred model was constructed using five feature extraction methods: Pseudo amino acid composition (PAAC), Composition transition distribution (CTDC), Dipeptide composition (DPC), Word2Vec, and FastText. Among these, FastText emerged as the most effective for encoding protein sequences. Despite the application of feature selection techniques, the analysis revealed that the original raw feature dimensions yielded superior results compared to the selected features. The model used seven machine learning classifiers, including Decision Tree (DT), Extra Trees Classifier (ETC), CatBoost (CAT), XGBoost (XGB), LightGBM (LGBM), Random Forest (RF), and the stacking model named SPPIPred. SPPIPred demonstrated exceptional accuracy rates of 0.9989 in the H pylori dataset and 0.9991 in the S cerevisiae dataset, with Matthews correlation coefficients (MCC) of 0.9982 and 0.9979, respectively. These findings highlight the effectiveness and reliability of the SPPIPred model, offering valuable insights to researchers in the field of bioinformatics and improving applications within bioengineering and pharmaceutical development.

Read PDF

Similar papers

Open access 2026

A Biologically Informed Hybrid Stacking Framework for Protein–Protein Interaction Prediction

Mapping the protein interactome is fundamental to understanding disease mechanisms and facilitating therapeutic development. Although protein language models (PLMs) such as ESM-2 have advanced protein-protein interaction (PPI) prediction, their high-dimensional representations remain difficult to connect to verifiable biological signals. To address this limitation, we propose HybridStack-PPI, a gray-box framework that combines ESM-2 sequence representations with explicit physicochemical and motif-derived biological descriptors. The architecture uses motif-anchored local pooling global mean pooling, symmetric pair encoding, fold-internal feature selection, LightGBM branch learners, and an elastic-net logistic-regression stacking layer. We evaluated the method using a C3 cluster-based cross-validation protocol with a 40% sequence-identity clustering threshold and a Same-GO hard-negative setting in which negative candidates shared functional annotations with positive pairs. Under this setting, HybridStack-PPI reached a Human ROC-AUC of 73.65%, PR-AUC of 91.35%, MCC of 28.06%, and specificity of 75.61%. The results indicate a conservative operating point: compared to more recall-oriented baselines, the proposed stack trades lower recall and F1 for higher specificity, MCC, and ranking behavior under functionally similar negative samples. We further reported cross-species transfer, ablation, latency, SHAP-based descriptor attribution, and meta-learner coefficient analyses to clarify both the promise and limitations of biologically informed PPI prediction.

T. T. Nguyen, X. Mai, N. Nguyen · 0 citations
Open access Aug 2026

3Dloop-FPSSM: Predicting Protein–Protein Interactions by Fusing 3D Local Optimal Oriented Pattern and Folded Position-Specific Scoring Matrix

Protein–protein interactions (PPIs) are fundamental to cellular processes, and understanding their mechanisms aids in disease diagnosis, drug target identification, and therapeutic development. Traditional experimental methods for PPI detection are costly and time-consuming, highlighting the need for efficient computational tools. In this study, we introduce a novel sequence-based framework for PPI prediction, which combines position-specific scoring matrices (PSSMs), 3D local optimal orientation patterns (3Dloop), and histogram gradient boosting (HistGB). Protein sequences are first transformed into PSSMs to capture evolutionary conservation, which are then processed into folded PSSMs (FPSSMs) to reveal hidden relationships among discontinuous amino acids. High-dimensional features are extracted using 3Dloop and classified with HistGB. We demonstrate the superiority of this method over random forest (RF) and support vector machine (SVM) models, achieving accuracies of 95.61% on the yeast dataset and 89.93% on the Helicobacter pylori dataset. Ablation studies confirm the effectiveness of each component in the framework. The results show that our approach provides a reliable and efficient solution for PPI prediction.

Fangfang Bai, Dan Liu, Guangxian Wang et al. · 0 citations
Open access Jul 2026

Predictions of protein–protein interactions: Learning sequences and structures

A neural network-based pipeline that integrates amino acid sequences with structural features is developed and provides a modular prototype for follow-up, more extensive protein modeling, including larger proteins and sequence of variable sizes.

Carl David Jasper Causin, M. Fyta · 0 citations
Open access Jul 2026

Explainable Machine Learning-Guided Virtual Screening and Molecular Docking for Identification of Novel FYN Kinase Inhibitors

FYN kinase is a non-receptor protein tyrosine kinase involved in various cancers and neurodegenerative diseases; however, no selective FYN inhibitor has been approved yet. Here we introduce the explainable Machine Learning (ML) coupled with virtual screening and Molecular Docking (MD) pipeline for fast prediction of new FYN kinase inhibitors. In this study, we constructed the training set of 906 molecules active against FYN kinase from the ChEMBL database. Molecules were encoded with Extended-Connectivity Fingerprints (ECFP4). The classification models Random Forest (RF) and eXtreme Gradient Boosting (XGBoost) were developed, and the latter showed the better performance in test (AUC=0.8118) and 5-fold cross-validation (AUC=0.8297). Based on the SHapley Additive exPlanations (SHAP) values obtained via TreeExplainer, nitrogen-containing heterocycles and hydrogen bond acceptors have been identified as the most important molecular substructures. Using the optimal XGBoost classifier, screening of 2,000 approved drugs has been performed, resulting in 470 hit molecules (23.5% hit rate). Five best molecules were further submitted to the MD procedure using AutoDock Vina to dock to FYN kinase domain (PDB RCSB: 2DQ7), showing binding energies in the interval of -9.57 to -6.32 kcal/mol. Dasatinib Anhydrous (CHEMBL1421) was the second strongest binder (-8.49 kcal/mol), effectively interacting with the ATP binding site. Although CHEMBL1171837 was the strongest binder (-9.57 kcal/mol), it was caught in the ADMET profiling. According to ADMET profiling, the top one inhibitor (CHEMBL1421) satisfies Lipinski’s rule of five and Veber rules. Analysis of hydrogen bond and hydrophobic interactions revealed hydrogen bonding with ASP148, LYS39, and ASN86 and hydrophobic interactions with ALA147, ILE80, and GLY88. Validation by self-docking procedure (self-docking or STS) showed low Root Mean Square Deviation (RMSD)<2.0 Å with a binding affinity of -11.53 kcal/mol. This work highlights how explainable ML can be used in combination with structure-based docking to expedite the drug discovery process against FYN kinase and can be applied to other kinase targets.

Ahmet Turan Demir · 4 citations