ClouDens is proposed, an anomaly detection framework tailored to LCS monitoring that leverages operational-context attributes encoded in the telemetry log schema to improve detection accuracy and early identification of anomalies.
Abstract
With the rapid growth of cloud computing infrastructures in scale and complexity, network monitoring for Large-scale Cloud Systems (LCSs) has become increasingly challenging, requiring automated and reliable anomaly detection to maintain service availability. Modern LCSs continuously generate telemetry logs from distributed cloud services, producing high-dimensional multivariate time series that capture system operations. Detecting anomalies in this context is difficult due to extreme dimensionality, complex dependencies among distributed components, and severe sparsity from intermittently active services. Taking these challenges into account, we first conduct an empirical study on telemetry logs from the IBM Cloud Console platform, and then propose ClouDens, an anomaly detection framework tailored to LCS monitoring that leverages operational-context attributes encoded in the telemetry log schema to improve detection accuracy and early identification of anomalies. ClouDens partitions high-dimensional telemetry logs into domain-guided subsets, constructs a context-aware graph modeling operational service dependencies, and employs Spatio-Temporal Graph Neural Networks for forecasting-based anomaly detection. We evaluate ClouDens on the recently released IBM Cloud Telemetry Dataset and provide practical insights into designing reliable anomaly detection solutions for LCS monitoring. Results show ClouDens achieves higher NAB scores in count-based telemetry features, indicating more accurate, earlier anomaly detection with broader coverage than a GRU-based model. Our study further reveals that telemetry feature subsets, operational-context modeling, scoring strategies, and sparsity imputation all substantially influence detection performance, offering practical guidance for designing and fairly benchmarking anomaly detection approaches for LCS monitoring.
This study created and verified an adaptive machine learning framework that makes use of real-time model updates and domain-specific cloud infrastructure information and offers a deployable framework for improving cloud security in Kenya and other resource-constrained environments.
Milimo Moses Sibilike· International Journal Of Eng...· 0 citations
Cloud services require continuous log monitoring to detect failures before outages occur. Most log anomaly detectors rely on fixed event identifiers or templates, which degrade as services evolve, parsers change, and new templates appear, leading to a mismatch between offline accuracy and online reliability. We propose...
System logs are critical for software reliability. While many automated log-based anomaly detection methods exist, they often falter in large-scale cloud systems due to high resource consumption and poor adaptability to evolving logs. In this paper, we present SeaLog, an accurate, lightweight, and adaptive log-based an...
Jinyang Liu, Junjie Huang, Zhihan Jiang et al.· ACM Transactions on Software...· 0 citations
A reliability-aware edge–cloud framework that treats early detection as a sequential routing problem, and identifies the minimum-evidence gate and cloud-refinement stage as the main reliability controls.
Siraj Azam, Farheen Naaz, Mikail Mohammed Salim· Electronics· 0 citations
Large-scale testing infrastructures are critical for validating telecommunication systems, yet their growing complexity makes efficient resource utilization and anomaly detection increasingly challenging. In reservation-based testbed environments, errors in resource allocation or preparation often manifest as abrupt sp...
Justyna Witulska, Marcin Szczukiewicz, Artur Tabaka et al.· 1 citation· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.