A cloud-native log analytics-based operating fault tracing method for power monitoring systems
Abstract
Reliable automated control in cloud-native power monitoring requires faults to be traced before stale measurements, delayed alarms, or unsafe commands propagate through optical-fiber Ethernet links and service chains. This study addresses the research problem of locating the initiating field, communication, or software component and selecting a safe recovery action under concurrent alarms and incomplete observability. TopoLog-FT first fuses raw-log semantics, service identity, synchronized metrics, distributed traces, and physical-to-digital topology into a relational temporal event graph. It then applies graph attention and an explicit causal-path head that extracts ancestor subgraphs, evaluates topology-consistent propagation paths, calibrates root-cause confidence, and gates restart, rollback, failover, queue throttling, fiber-path rerouting, or operator escalation through risk constraints. Experiments on a 12-node Kubernetes-based testbed with 18 mi- croservices, 1.82 million log records, and 320 controlled injections of eight fault types yielded 92.3% Top-1 root-cause accuracy, 98.1% Top-3 accuracy, 94.6% Macro-F1, and 0.68 s mean tracing latency at 10,000 events/s. Relative to the strongest baseline, the method improved Top-1 accuracy by 4.8 percentage points, reduced mean time to recovery by 38.7%, and limited unsafe automated actions to 2.5%. Topology-aware log analytics therefore enables dependable closed-loop automation while strengthening the continuity and diagnosability of optical communication paths in power monitoring systems.