A Distributed Observability Framework for Predictive Reliability and Failure Prevention in Large-Scale Cloud Infrastructure
Aim: This study aimed to develop, implement, and evaluate a Distributed Observability Framework (DOF) for proactive reliability management in large-scale cloud-native infrastructures. The framework integrates telemetry collection, dependency-aware observability analysis, machine learning, dynamic observability scoring,...