When Calibrated Detectors Meet New Attacks: Per-Category Reliability of Machine-Learning Intrusion Detection Under Distribution Shift
Machine-learning intrusion detectors are usually reported with accuracy or F1 on a single train and test split, and their confidence scores are often read operationally as probabilities without an explicit calibration check. We test that assumption. We measure the reliability of the predicted probabilities of three cla...