ST-HAE detects spatio-temporal anomalies in heterogeneous IoT networks under realistic class imbalance
TL;DR
Accuracy on IoTMal-2026, matching or exceeding Deep SVDD, Deep SAD, and Kitsune under an identical protocol is evaluated, finding the model’s bidirectional recurrent component does not reliably improve mean accuracy over a simpler, convolution-only alternative.
Abstract
Most intrusion detection systems (IDS) research is conducted on artificially balanced datasets. This hides a real problem: precision drops sharply under production-like conditions, where attacks can be outnumbered by benign packets hundreds to one. This paper evaluates an unsupervised detector under exactly that kind of imbalance, on two structurally different datasets: consumer IoT traffic (CIC-IoT 2023, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$278.5\times$$\end{document} attack-to-benign) and multi-architecture malware traffic (IoTMal-2026, \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$9.6\times$$\end{document}). The detector, ST-HAE (Spatio-Temporal Hybrid Autoencoder), is a 1D-CNN-BiLSTM encoder-decoder with 60,151 parameters. It trains only on benign traffic and flags anomalies using a statistically calibrated reconstruction-error threshold, with no access to attack labels at any point. Under five-seed statistical validation, ST-HAE reaches alert precision \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$0.9999\!\pm \!0.0000$$\end{document} on CIC-IoT 2023 and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$95.85\,\%\!\pm \!0.03\,\%$$\end{document} accuracy on IoTMal-2026, matching or exceeding Deep SVDD, Deep SAD, and Kitsune under an identical protocol. Deep SVDD is also notably unstable, with accuracy standard deviation up to 10.0 percentage points across seeds; ST-HAE and Kitsune do not share this problem. Separately, we find the model’s bidirectional recurrent component does not reliably improve mean accuracy over a simpler, convolution-only alternative. What it does provide is stability: training-run failure risk drops roughly 31-fold on the more heterogeneous dataset, and we treat that stability, not raw accuracy, as its real contribution. A category-level breakdown tells a more complicated story about precision. Recall on high-volume flood attacks (DDoS, DoS, Mirai) is near-perfect. Recall on slower, quieter attacks (reconnaissance, web exploits, spoofing) is far worse, with reconstruction error running up to \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$454\times$$\end{document} smaller for these attacks. Threshold sweeps, using both Gaussian and nonparametric calibration, confirm this is a real separation limit rather than something a different threshold could fix—which matches what the paper’s decision-theoretic threshold model predicts once correctly re-derived in this revision to account for class separation: below a critical separation point, the theoretically optimal threshold simply is not usable in practice. Feature attribution, cross-checked with SHAP and LIME, shows flood detection depends heavily on one or two count-based features that exceed the training data’s normal range. Removing those features does not eliminate detection, but the effect is uneven: Mirai detection is barely affected, while DoS recall drops from near-total to \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$59\,\%$$\end{document}. This rules out both extremes—the features are not the whole story, but they are not irrelevant either. Finally, we report single-sample inference latency measured directly on real AWS Graviton ARM64 hardware: 1.82 ms per sample, 550 samples/s, correcting an earlier estimate that had been extrapolated from GPU numbers rather than measured. ST-HAE is best understood not as a universally strong detector, but as a precision-focused, edge-deployable one whose real capabilities and real limits are both measured here, under realistic and imbalanced conditions.