Skip to content
Open access

ST-HAE detects spatio-temporal anomalies in heterogeneous IoT networks under realistic class imbalance

Aug 2026 · Discover Networks · Vol 2 · 0 citations · 23 references

TL;DR

Accuracy on IoTMal-2026, matching or exceeding Deep SVDD, Deep SAD, and Kitsune under an identical protocol is evaluated, finding the model’s bidirectional recurrent component does not reliably improve mean accuracy over a simpler, convolution-only alternative.

Abstract

Most intrusion detection systems (IDS) research is conducted on artificially balanced datasets. This hides a real problem: precision drops sharply under production-like conditions, where attacks can be outnumbered by benign packets hundreds to one. This paper evaluates an unsupervised detector under exactly that kind of imbalance, on two structurally different datasets: consumer IoT traffic (CIC-IoT 2023, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$278.5\times$$\end{document} attack-to-benign) and multi-architecture malware traffic (IoTMal-2026, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$9.6\times$$\end{document}). The detector, ST-HAE (Spatio-Temporal Hybrid Autoencoder), is a 1D-CNN-BiLSTM encoder-decoder with 60,151 parameters. It trains only on benign traffic and flags anomalies using a statistically calibrated reconstruction-error threshold, with no access to attack labels at any point. Under five-seed statistical validation, ST-HAE reaches alert precision \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$0.9999\!\pm \!0.0000$$\end{document} on CIC-IoT 2023 and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$95.85\,\%\!\pm \!0.03\,\%$$\end{document} accuracy on IoTMal-2026, matching or exceeding Deep SVDD, Deep SAD, and Kitsune under an identical protocol. Deep SVDD is also notably unstable, with accuracy standard deviation up to 10.0 percentage points across seeds; ST-HAE and Kitsune do not share this problem. Separately, we find the model’s bidirectional recurrent component does not reliably improve mean accuracy over a simpler, convolution-only alternative. What it does provide is stability: training-run failure risk drops roughly 31-fold on the more heterogeneous dataset, and we treat that stability, not raw accuracy, as its real contribution. A category-level breakdown tells a more complicated story about precision. Recall on high-volume flood attacks (DDoS, DoS, Mirai) is near-perfect. Recall on slower, quieter attacks (reconnaissance, web exploits, spoofing) is far worse, with reconstruction error running up to \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$454\times$$\end{document} smaller for these attacks. Threshold sweeps, using both Gaussian and nonparametric calibration, confirm this is a real separation limit rather than something a different threshold could fix—which matches what the paper’s decision-theoretic threshold model predicts once correctly re-derived in this revision to account for class separation: below a critical separation point, the theoretically optimal threshold simply is not usable in practice. Feature attribution, cross-checked with SHAP and LIME, shows flood detection depends heavily on one or two count-based features that exceed the training data’s normal range. Removing those features does not eliminate detection, but the effect is uneven: Mirai detection is barely affected, while DoS recall drops from near-total to \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$59\,\%$$\end{document}. This rules out both extremes—the features are not the whole story, but they are not irrelevant either. Finally, we report single-sample inference latency measured directly on real AWS Graviton ARM64 hardware: 1.82 ms per sample, 550 samples/s, correcting an earlier estimate that had been extrapolated from GPU numbers rather than measured. ST-HAE is best understood not as a universally strong detector, but as a precision-focused, edge-deployable one whose real capabilities and real limits are both measured here, under realistic and imbalanced conditions.

Read PDF

Similar papers

Open access Aug 2026

Statistical Drift detection with parametric Gamma–Weibull and adaptive deep learning for malware detection in IoT data streams

Machine learning (ML) and deep learning (DL) models for Internet of Things (IoT) malware detection may experience performance degradation in non-stationary data streams due to evolving malware behavior, system updates, and dynamic network conditions. In this study, we present a statistically based drift-adaptive framew...

F. Agu, Robert Andok, Jaromír Klarák et al. · 0 citations
#edge computing Sep 2026

An experimental study of edge deployment strategies for cryptocurrency analytics workloads

An empirical study of two production edge platforms using latency-sensitive cryptocurrency analytics workloads, conducted from a local Southeast-Asian client and three supplemental AWS cloud clients, reveals that platform-level latency advantages are dominated by client-edge proximity rather than systemic runtime diffe...

M. Tang · 0 citations
Open access Aug 2026

Strong Hotspot Domination Integrity of Corona Products, Comb Products and Shadow Graphs with Applications

The study of graph vulnerability parameters provides valuable insights into the structural stability and resilience of networks when there is a disruption. The strong hotspot domination integrity (DISH\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \use...

A. Antony, V. Sangeetha · 0 citations

Explainable Federated Learning for Trustworthy Thoracic Disease Detection Under Non-IID Data Distributions

Comparison of SHAP and Grad-CAM attribution maps confirms clinically coherent disease-specific localisation, and reveals monotonic performance degradation, identifying minimal regularisation as optimal for multi-label medical imaging.

S. Arumugam, A. Sindhu, M. N. Saroja et al. · 0 citations
Open access Aug 2026

One (Noisy) Bit to Rule Them All: Key Recovery from Randomness Leakage in ML-DSA

The Fiat-Shamir transform is one of the most widely applied methods for secure signature construction. Fiat-Shamir starts with an interactive zero-knowledge identification protocol and transforms this via a hash function into a non-interactive signature. The protocol’s zero-knowledge property ensures that a signature d...

Simon Damm, Nicolai Kraus, Alexander May et al. · 1 citation · ⚡1
#graph neural networks Open access Sep 2026

Edge weight concentration overcomes node degree blindness in graph based network intrusion detection

Graph-based network intrusion detection almost universally represents attacker behavior through node-centric structural features – degree, centrality, embeddings from graph neural networks – which implicitly assume that an attacker’s identity is visible as a distinct node in the host-communication graph. We show that t...

Md Hasibuzzaman · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.