Skip to content
Open access

When Calibrated Detectors Meet New Attacks: Per-Category Reliability of Machine-Learning Intrusion Detection Under Distribution Shift

2026 · International Journal of Advanced Computer Science and Applications · 0 citations · 45 references

Abstract

Machine-learning intrusion detectors are usually reported with accuracy or F1 on a single train and test split, and their confidence scores are often read operationally as probabilities without an explicit calibration check. We test that assumption. We measure the reliability of the predicted probabilities of three classifiers, logistic regression, random forest, and XGBoost, on two benchmarks, NSL-KDD and UNSW-NB15, in two settings: when training and test data share a distribution, and under the shift each benchmark's official split already contains. In distribution, all three detectors are well calibrated, with top-label expected calibration error at most 0.03. Under shift, the outcome depends on its content. On NSL-KDD, whose test set contains attack types absent from training, top-label calibration error rises from at most 0.007 in distribution to between 0.170 and 0.209, and the error is largest on the benign and Remote-to-Local classes, so unfamiliar attacks are labeled normal with high confidence. On UNSW-NB15, whose split keeps the same attack categories, the error stays far smaller, between 0.025 and 0.068, even though a domain classifier detects a clear shift on both benchmarks. Recalibrating on training-distribution data does not reliably transfer to the shifted test set. Supervised target-domain recalibration on a held-out labeled sample from the shifted distribution substantially reduces the calibration error and recovers much of the minority-class recall on NSL-KDD. Most of the aggregate top-label ECE reduction is obtained with about one hundred labeled target samples. This remedy requires labeled observations from the shifted environment and therefore describes reliability after target labels become available, not reliability against attacks that remain entirely unseen. A decision-curve analysis and expected-cost comparison show that an uncalibrated score used at its nominal cost threshold is very costly, and that both recalibration and a target-tuned threshold reduce operating cost, with recalibration additionally improving the multiclass probabilities. A leave-one-attack-out study reproduces the failure and indicates that the attack family for which performance degrades substantially is the one least separable from benign traffic, not the one most distinct from the other attacks.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.