Calibration Collapse in Federated Industrial IoT Intrusion Detection Under Non-IID Heterogeneity
Abstract
Federated learning (FL) has been proposed for privacy-preserving Industrial Internet of Things (IIoT) intrusion detection, and predictive uncertainty is expected to support zeroday attack recognition. We compare centralized learning with FedAvg, Mean, Trimmed Mean, Krum, and DP-FedAvg on nearindependent and identically distributed (IID) synthetic Edge-IIoT and highly non-IID TON_IoT datasets. Results show that FL remains viable with adequate training budgets, achieving performance close to centralized learning on near-IID data and an F1 Score of 0.91 on TON_IoT. However, calibration degrades substantially under FL, with expected calibration error increasing by 2-9× despite competitive F1-scores, resulting in severe recall-false-alarm tradeoffs. Differential privacy further degrades performance, while Byzantine-resilient aggregators fail under strong non-IID conditions. The primary factor is non-IID severity, with calibration collapse emerging as the common failure mechanism. These findings suggest that IIoT federated IDS evaluation should extend beyond accuracy and F1-score to include calibration and open-set robustness, and that calibrationaware thresholding may be more critical than developing new aggregation methods.