Skip to content
Open access

Reliable Sequenced-Based Protein-Protein Interaction Prediction Using Lempel Ziv Complexity and Optimized Deep Learning Model

2026 · IEEE Access · Vol 14, pp. 103756-103775 · 0 citations · 60 references
Computer Science

TL;DR

The novel combination of LZ complexity–based negative sample selection, CT feature representation, and GA-optimized CNN–LSTM architecture provides a robust and biologically informed framework for PPI prediction.

Abstract

In this study, a hybrid Convolutional Neural Network–Long Short-Term Memory (CNN–LSTM) deep learning model was developed for sequence-based prediction of protein–protein interactions (PPIs). To address the limitations of experimental methods in terms of time and cost, computational approaches were employed. A novel method based on Lempel–Ziv (LZ) complexity was proposed to select reliable non-interacting protein pairs. Protein sequences were represented as feature vectors using the Conjoint Triad (CT) method, which encodes amino acid physicochemical properties. The hybrid CNN–LSTM architecture was then used to classify interacting and non-interacting protein pairs, where CNN layers captured local sequence motifs and LSTM layers modeled long-range dependencies. Furthermore, Genetic Algorithm (GA)–based hyperparameter optimization was applied to tune model hyperparameters. The novelty of this study lies in the combination of LZ complexity–based negative sample selection, CT feature representation, and GA-optimized CNN–LSTM architecture, providing a robust and biologically informed framework for PPI prediction. The proposed model achieved 91% training accuracy and 90% testing accuracy before optimization, which increased to 94% and 93%, respectively, after GA optimization. These results demonstrate that the integrated approach enhances predictive performance and enables reliable extraction of meaningful information from protein sequences.

Read PDF

Similar papers

Open access 2026

A Biologically Informed Hybrid Stacking Framework for Protein–Protein Interaction Prediction

Mapping the protein interactome is fundamental to understanding disease mechanisms and facilitating therapeutic development. Although protein language models (PLMs) such as ESM-2 have advanced protein-protein interaction (PPI) prediction, their high-dimensional representations remain difficult to connect to verifiable biological signals. To address this limitation, we propose HybridStack-PPI, a gray-box framework that combines ESM-2 sequence representations with explicit physicochemical and motif-derived biological descriptors. The architecture uses motif-anchored local pooling global mean pooling, symmetric pair encoding, fold-internal feature selection, LightGBM branch learners, and an elastic-net logistic-regression stacking layer. We evaluated the method using a C3 cluster-based cross-validation protocol with a 40% sequence-identity clustering threshold and a Same-GO hard-negative setting in which negative candidates shared functional annotations with positive pairs. Under this setting, HybridStack-PPI reached a Human ROC-AUC of 73.65%, PR-AUC of 91.35%, MCC of 28.06%, and specificity of 75.61%. The results indicate a conservative operating point: compared to more recall-oriented baselines, the proposed stack trades lower recall and F1 for higher specificity, MCC, and ranking behavior under functionally similar negative samples. We further reported cross-species transfer, ablation, latency, SHAP-based descriptor attribution, and meta-learner coefficient analyses to clarify both the promise and limitations of biologically informed PPI prediction.

T. T. Nguyen, X. Mai, N. Nguyen · 0 citations
Open access Aug 2026

DHST: A Deep Hybrid Structure–Topology Framework for Accurate Protein Function Prediction

Accurate protein function prediction (PFP) is essential for understanding biological systems. However, structure-based graph neural networks often rely on fixed-distance contact maps, which may inadequately capture continuous, multi-scale spatial topologies, while the long-tail distribution of Gene Ontology (GO) labels may bias prediction toward frequent functions. We propose DHST, a deep hybrid structure–topology framework that integrates sequence semantics from a pretrained protein language model with local structural information learned by a residual graph convolutional network. DHST further introduces site-specific persistent homology to encode multi-scale topological invariants and a topology-guided residue-wise gated fusion module to modulate structure–semantics representations using local topological embeddings. The fused residue features are aggregated through dual-path pooling, and a weighted binary cross-entropy loss is used to mitigate the adverse effects of label imbalance. On the PDB dataset, DHST achieved area under the precision–recall curve (AUPR) scores of 0.779, 0.481, and 0.557 for molecular function (MF), biological process (BP), and cellular component (CC), respectively; on the AF2 dataset, the corresponding scores were 0.729, 0.390, and 0.459. The model also demonstrated robust generalization to low-homology proteins and maintained strong predictive performance across GO terms with different levels of functional specificity. Ablation results supported the contributions of the main components.

Bin Lu, Fujun Xiang, Hai-Long Wang et al. · 0 citations
Open access Jul 2026

SPPIPred: Stacking-based ensemble learning model for identification of protein-protein interaction

SPPIPred, an advanced machine learning-based model designed for precise PPI prediction, is presented, offering valuable insights to researchers in the field of bioinformatics and improving applications within bioengineering and pharmaceutical development.

M. Rahman, M. Ali, Md. Shohidullah et al. · 0 citations
Open access Jul 2026

Predictions of protein–protein interactions: Learning sequences and structures

A neural network-based pipeline that integrates amino acid sequences with structural features is developed and provides a modular prototype for follow-up, more extensive protein modeling, including larger proteins and sequence of variable sizes.

Carl David Jasper Causin, M. Fyta · 0 citations
Open access Aug 2026

DeepPNI: a language- and graph-based model for mutation-driven protein–nucleic acid binding energetics

DeepPNI is a deep learning regression model that integrates sequence- and structure-based features to estimate mutation-induced changes in binding free energy in protein–nucleic acid complexes, developed using a comprehensive dataset of 1754 mutations spanning protein–DNA and protein–RNA complexes.

Somnath Mondal, Tinkal Mondal, Soumajit Pramanik et al. · 0 citations