The proposed Attention-Augmented CNN-BiLSTM architecture incorporating a lightweight Temporal-Spatial Attention Module (TSAM) achieves an effective balance between computational efficiency, robustness to severe class imbalance, and accurate multi-class intrusion detection for next-generation IoT security systems.
Abstract
Intrusion detection in heterogeneous network environments remains challenging because of severe class imbalance, evolving attack patterns, and the limited ability of existing deep learning models to jointly exploit temporal dependencies and feature-level correlations. To address these challenges, this paper proposes an Attention-Augmented CNN-BiLSTM (AA-CNN-BiLSTM) architecture incorporating a lightweight Temporal-Spatial Attention Module (TSAM) that independently models temporal and feature-wise attention through factorised attention streams and adaptively fuses them using a learnable gating mechanism. Unlike conventional self-attention, TSAM achieves a computational complexity of O(TC(T + C)), making it suitable for resource-constrained edge internet-of-things (IoT) environments. Furthermore, severe class imbalance is mitigated through the combined use of SMOTE oversampling and Focal Loss, providing complementary data-level and objective-level imbalance handling. The proposed framework is evaluated on two publicly available benchmarks representing different network settings: CIC-IoV-2024, which provides a vehicular-IoT intrusion detection environment, and CIC-IDS2017, which serves as a general network intrusion detection benchmark. On the highly imbalanced CIC-IoV-2024 dataset, the proposed model achieves a Macro-F1 score of 0.7421, Macro-AUC of 0.9961, and Cohen’s Kappa of 0.9374, representing improvements of 2.81 percentage points in Macro-F1 and 0.0167 in Kappa over the CNN-BiLSTM baseline while maintaining an overall accuracy of 98.71%. On CIC-IDS2017, the proposed framework attains a Macro-F1 score of 0.7951, Macro-AUC of 0.9903, and Cohen’s Kappa of 0.8841, outperforming the baseline by 2.01 percentage points in Macro-F1 and 0.0249 in Kappa. Additional ablation studies confirm the complementary contributions of SMOTE and Focal Loss, while inference latency analysis indicates that TSAM introduces only a modest computational overhead and preserves millisecond-level inference suitable for real-time deployment. Statistical significance of the observed improvements is confirmed using paired Wilcoxon signed-rank tests across six independent experimental runs. Overall, the proposed AA-CNN-BiLSTM achieves an effective balance between computational efficiency, robustness to severe class imbalance, and accurate multi-class intrusion detection for next-generation IoT security systems.
Detecting cyber intrusions in modern IoT networks is challenging because of their large scale, heterogeneous device ecosystems, and high-volume traffic patterns. This paper presents a cross-attention CNN–LSTM fusion architecture that jointly learns the spatial and temporal characteristics of network traffic for bin...
Mohamed Fakri, A. Najid, Rachid Ben Said et al.· Scientific Reports· 0 citations
With the rapid advancement of network technology and the Internet of Things (IoT), massive, high-dimensional traffic data pose significant challenges to Network Intrusion Detection Systems (NIDS). Existing deep learning methods face two major limitations: (1) Insufficient feature extraction: single models struggle to c...
Ting-Hui Huang, Yu Wang, Yu-Ming Qin· Italian National Conference...· 0 citations
Distributed Denial of Service (DDoS) attacks continue to pose a severe and escalating threat to networked digital infrastructure, with global attack volumes rising by over 53% in 2024 alone. While machine learning approaches have demonstrated improved detection performance over traditional rule-based systems, many exis...
Godson Samwel, I. Tende, Gustaph Sanga· East African Journal of Info...· 0 citations
The Tor network’s anonymity is increasingly exploited for cybercrime, creating a demand for accurate traffic classification under strict few-shot constraints. While recent efforts like WF-Transformer demonstrate strong temporal modeling capabilities, they still require abundant labeled data and struggle to generalize u...
De-Peng Xu, Guo-Zhen Cheng, Hong-Chao Hu et al.· IEEE Transactions on Network...· 0 citations
The principal contribution of this work is architectural and diagnostic rather than a performance improvement: it documents that combining feature-wise attention with out-of-fold stacked generalization does not, in this setting, outperform a plain multi-layer perceptron, while incurring the highest memory footprint of...
Mahima Khanna, V. Murthy, Siva Ramavarapu et al.· International Journal for Gl...· 0 citations
One lesson emerges from the experiments: putting the effort into how traffic is written down, instead of making the classifier heavier, offers an economical and workable path to intrusion detection across heterogeneous network environments.
Asmaa Benchama, Khalid Zebbara· EPJ Web of Conferences· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.