CTGuard: A Calibrated Cross-System Log Anomaly Detection System for Resilient Cloud and HPC Operations
Abstract
While system logs can be used to alert cloud and high-performance computing operators to incidents, real-world log anomaly detectors are often unreliable when log streams vary across platforms, when anomalies are highly imbalanced, or when alert thresholds are chosen without proper operational risk control. In this work, we introduce CTGuard, a lightweight calibrated log anomaly detection system designed for heterogeneous infrastructure logs. CTGuard converts logs into sequences of templates and severity tokens, constructs class-conditional template and transition probabilities, and tunes the final alert threshold using a validation split under an explicit false-positive budget. The system is deployment ready with no need for new dataset collection, no large language model inference, and provides interpretable scores through event and transition-level contributions. We evaluate CTGuard on two public LogHub datasets; HDFS and BGL, representing distributed cloud storage and high-performance computing environments. CTGuard achieves F1 scores of 0.937 on HDFS and 0.740 on BGL, outperforming frequency likelihood, Markov likelihood, and an uncalibrated template classifier. It maintains false positive rates below 0.3% on HDFS and below 5% on BGL, with inference time under 0.008 ms per test sequence in our Node.js implementation. The results show that statistically grounded, calibrated approaches provide a strong and practical foundation for resilient system operations.