Adapting and Validating GNN Fault Forecaster with Real-World Cloud Traces
Abstract
Modern cloud applications are built using microservices that support various sectors such as healthcare, finance, ecommerce and communication. Microservices’ distributed and scalable nature provides high performance and flexibility but also increases system complexity. When failures occur, they propagate rapidly through interconnected microservices, leading to performance degradation, service downtime and financial losses. Traditional fault detection approaches mainly depend on threshold-based rules and reactive fault tracing, which are time-consuming, difficult to scale and often fail to proactively detect faults in dynamic service environments. Existing research has proposed a Graph Neural Network (GNN)-based framework that models microservice interactions as graphs for fault prediction and critical component identification. The framework, trained on simulated data generated using Markov Decision Processes (MDPs), has demonstrated strong performance in controlled environments. However, its use of synthetic data limits its effectiveness on real-world cloud systems, which exhibit unpredictable behaviour and dynamic workloads. The proposed work addresses this limitation using the Alibaba Cluster Trace – Microservices v2021 dataset collected from a real production environment. This work incorporates an Edge-Attention Graph Convolutional Network (EA-GCN), a hybrid loss function and Critical Node Removal Analysis (CNRA) to improve fault prediction and identify critical microservices. Experimental results demonstrate that the proposed approach achieves an F1-score of 0.7265 and an AUC-ROC of 0.9583, improving F1-score by 47.9% over plain GCN. CNRA identifies the most critical microservices that contribute to system-wide fault risk, prioritizing fault mitigation. The results show that the framework is effective for fault prediction and critical microservice identification in real-world cloud environments.