Skip to content
Open access

Feature-Space Adversarial Robustness Evaluation of Lightweight ML-Based Intrusion Detection for IoT and Healthcare-Sensing Networks

2026 · IEEE Access · Vol 14, pp. 141428-141454 · 0 citations · 55 references

Abstract

Machine learning (ML)-based Network Intrusion Detection Systems (NIDS) are increasingly deployed in IoT and healthcare-sensing networks. Reported robustness depends on how it is measured: evaluations relying on attacks transferred from a surrogate cannot separate a robust decision boundary from a merely dissimilar one, and computational cost is rarely reported alongside accuracy. This paper contributes an evaluation protocol rather than a new attack or defense. Four primary lightweight classifiers (logistic regression, linear SVM, random forest, and a compact MLP), together with two supplementary gradient-boosted targets (XGBoost and LightGBM), are assessed on CSE-CIC-IDS2018, TON_IoT, and WUSTL-EHMS-2020, a healthcare benchmark pairing network-flow features with patient biometrics. Surrogate-based gradient attacks (FGSM, PGD, C&W) are paired with a direct decision-based HopSkipJump baseline and with gradient attacks computed on the linear models themselves, separating transfer effects from direct vulnerability. Training-time hardening, adversarial training for the MLP and surrogate-augmented retraining for the classical models, is compared with randomized smoothing. All perturbations are generated in standardized feature space, so reported values are upper bounds on realizable evasion, not executable packet-level attacks. A complementary evaluation freezes protocol- and physiologically constrained features to bound that gap. Every attack family raises the false-negative rate above 0.78 on multiple model–dataset pairs, spanning three orders of magnitude where the clean rate is near zero. Classical models retain higher accuracy under transferred attacks but drop sharply under direct attack, so transfer-only evaluation overestimates their robustness. Adversarial training yields more consistent gains than randomized smoothing at greater training cost. Healthcare-sensing findings rest on a single benchmark and are reported as such. Robustness, transferability, false-negative rate, and inference cost need to be evaluated jointly.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.