LSTM-PPO: A Hybrid Deep Reinforcement Learning Framework with Asymmetric Reward Engineering for Real-Time Intrusion Detection in IIoT and SCADA Networks
Aug 2026· 2026 6th International Conference on Emerging Smart Technologies and Applications (eSmarTA)· pp. 1-7· 0 citations· 24 references
Abstract
The integration of SCADA systems with the Industrial Internet of Things (IIoT) has dramatically expanded the attack surface of critical infrastructure. Traditional intrusion detection systems (IDS) struggle with evolving threats, class imbalance (attack samples often <5%), and real-time constraints (<100 ms). Deep reinforcement learning (DRL) offers a sequential decision-making paradigm that adapts over time. This paper presents a hybrid LSTM-PPO framework that unifies temporal feature extraction (LSTM), synthetic minority oversampling (SMOTE), asymmetric reward engineering, and Proximal Policy Optimization (PPO). The LSTM captures multi-stage attack patterns, SMOTE addresses class imbalance exclusively on training data to prevent leakage, and the asymmetric reward heavily penalizes false negatives (-50) compared to false positives (-10), aligning with industrial safety priorities. PPO ensures stable and efficient policy learning. Extensive experiments on three benchmark datasets (WUSTL-IIoT-2021, NF-UNSW-NB15-v2, WUSTL-SCADA-2018) demonstrate near-perfect detection (up to 100% F1 on WUSTL-IIoT-2021, 99.99% accuracy on NF-UNSW-NB15-v2, 99.96% on WUSTL-SCADA-2018) with sub-microsecond inference latency (≈1 μs per sample on GPU batching, <25 μs for single sample). Cross-validation and ablation studies confirm robustness against overfitting and the contribution of each component. The framework meets real-time industrial requirements and outperforms state-of-the-art supervised and DRL-based IDS. Limitations include binary classification and adversarial robustness, which are left for future work.
Real-world IoT network security generates traffic at big-data scale with extreme class imbalance, temporal non-stationarity, and continuously evolving attack strategies that overwhelm static supervised classifiers. This paper presents a cognitive computing framework for network intrusion detection: a CNN–LSTM–DQN archi...
Xin Su, Zhiquan Bai, K. Ramli et al.· Big Data and Cognitive Compu...· 1 citation
Distributed Denial-of-Service (DDoS) attacks remain among the most disruptive network threats, and detectors that generalize across attack families with low false-alarm rates are still an open problem. Propose an adaptive hybrid ensemble that unifies two gradient-boosting learners (Random Forest and Gradient Boosting)...
Unknown authors· International Journal of Adv...· 0 citations
ShieldDRLNet is a hybrid deep reinforcement learning framework for proactive cloud-network intrusion detection that employs a convolutional neural network and a long short-term memory encoder to obtain a spatiotemporal traffic representation and uses a Double Deep Q-Network agent for adaptive sequential decision-making...
S. Venkatramulu, Anitha Patil, K. Pradeep et al.· Discover Computing· 0 citations
A Large Language Model-enhanced Autonomous Reinforcement Learning Penetration Testing framework that leverages the domain knowledge embedded in a Large Language Model to perform tactical planning, thereby pruning the original action space into a compact set of candidate actions.
A D2ANN-RL framework that integrates input/output sanitization, context isolation, sandboxing, and secure prompt engineering, supported by hybridization of Artificial Neural Network (ANN)–Reinforcement Learning (RL) detection model is introduced.
Victor Omoboye Oluwasegun, O. Falebita, Nabeela Temitayo Adebola et al.· Scientific Journal of Comput...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.