Skip to content
Open access

The Aggregate–Worst-Class Trade-Off in Feature Selection for IoT Intrusion Detection: A Multi-Objective Study on TON_IoT

Oct 2026 · Computation · 0 citations · 14 references

Abstract

Machine-learning intrusion detection on multi-attack Internet-of-Things (IoT) datasets is often reported through aggregate metrics that can hide near-failure on rare attack classes, and through single hand-picked feature subsets whose stability is seldom examined. Using the real TON_IoT Network dataset (211,043 flows, 10 classes, with man-in-the-middle (MITM) traffic at 0.5% prevalence), this study contributes a transparent empirical mapping rather than a new detector. A strong full-feature Random Forest reaches 0.969 accuracy and 0.945 macro-averaged F1 (macro-F1) yet recalls only 0.776 of the rare MITM class, and a mutual-information filter raises aggregate macro-F1 to 0.948 while leaving worst-class recall unchanged at 0.776. A self-coded non-dominated sorting genetic algorithm (NSGA-II) is then used to map the trade-off between macro-F1 and worst-class recall as the selected subset varies, alongside a matched-size random-subset control, standard imbalance remedies, and two additional classifiers. Three findings are robust. First, imbalance remedies on the full feature set—class weighting and, especially, the synthetic minority over-sampling technique (SMOTE)—raise MITM recall to 0.818 and 0.904, matching or exceeding the wrapper, so feature selection is not the most effective tool for the rare class. Second, re-adding a single destination-port feature inflates ransomware recall by 0.094 and cross-site-scripting (XSS) recall by 0.066 while doing nothing for MITM, showing that several apparently detected classes are separable by an identifier artifact whereas the genuinely hard class is not. Third, selected subsets are unstable (Nogueira index 0.21) despite reproducible objective values. The wrapper’s own rare-class advantage is small and does not reach significance under a cross-validation-appropriate corrected test. The two transferable findings replicate on the NSL-KDD benchmark. We argue that IoT intrusion-detection studies should report the aggregate–worst-class frontier, an identifier ablation, a selection-stability index, and cross-validation-appropriate significance tests.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.