Skip to content
Review Open access

The Future of Site Reliability Engineering: AI-Driven Observability and Autonomous Operations in Multi-Cloud Environments

2024 · International Journal of Emerging Trends in Computer Science and Information Technology · Vol 5, pp. 204-214 · 0 citations

TL;DR

The results indicate that AI and ML in observability systems can deliver automated operational workflows that help enable proactive reliability management, minimize downtime and accelerate decision making, which makes AIOps powered autonomous operations a rapidly growing important enabler for cloud native infrastructures with self repairing systems.

Abstract

SRE has come a long way since its start, from monitoring infrastructure health to reactive events. Modern SRE is an AI-driven observability platform providing real-time visibility into complex distributed systems. As multi-cloud use grows, operational complexity increases and it becomes increasingly complicated to provide dependability, performance and security across a broad range of cloud platforms. Artificial Intelligence (AI), Machine Learning (ML) and AIOps technologies are changing Service Availability and Incident Response (SRE) with intelligent anomaly detection, predictive analytics, automated root-cause investigation and autonomous remediation in response to such. The effort aims at investigating the future of service-oriented architecture (SRE) in the age of AI-based observability and autonomous operations in multi-cloud environments. Through a review of current technologies, industry practice and upcoming trends in research. The objective of this research is to investigate the viability of the application of AI-based solutions to enhance system dependability, minimize operational overhead and optimize incident response efficiency. The study will also address issues of scalability, interoperability and governance. The results indicate that AI and ML in observability systems can deliver automated operational workflows that help enable proactive reliability management, minimize downtime and accelerate decision making. This makes AIOps powered autonomous operations a rapidly growing important enabler for cloud native infrastructures with self repairing systems. This paper describes the convergence of AI with software defined networking (SRE) and strategic implications for enterprises seeking strong, scalable and efficient cloud operations. The results show intelligent automation is becoming more important in shaping the future of cloud-native reliability management and operational excellence.

Read PDF

Similar papers

2026

Data-driven fault detection and diagnosis in automated manufacturing systems

Data-driven methods, in contrast to traditional reactive ones, make use of continuous sensor data streams to anticipate anomalies before they can result in critical failures using advanced analytics, thereby improving responsiveness and efficiency in the system.

V. K. Nassa · 0 citations
#federated learning Open access Aug 2026

LACSF framework deployment for fault detection and mitigation

AI-driven failure detection is becoming essential in industrial manufacturing systems where conventional diagnostic methods often fall short in reliability and live feedback. This paper presents the integration of a modular artificial intelligence framework adapted to overcome these challenges by enabling intelligent f...

Faisal Shaikh, Sudipt Panta, R. R. Kumar et al. · 0 citations
Review Open access Aug 2026

From Detection to Verified Action: Operational Readiness for AI-Enabled Cloud Failure Management

The proposed Operational Decision-Readiness and Verification framework represents a deployment claim through analytical capability, operational authority, and assurance maturity and is conceptual and requires prospective and inter-rater validation.

Adepegba Akindayomi Akintade · 0 citations
Open access Sep 2026

A Distributed Observability Framework for Predictive Reliability and Failure Prevention in Large-Scale Cloud Infrastructure

Aim: This study aimed to develop, implement, and evaluate a Distributed Observability Framework (DOF) for proactive reliability management in large-scale cloud-native infrastructures. The framework integrates telemetry collection, dependency-aware observability analysis, machine learning, dynamic observability scoring,...

Chiranjevi Sai Venkat Kaushik Dhulipudi · 0 citations
Open access Sep 2026

A rule-driven SOC architecture for real-time threat detection and autonomous incident response

A rule-driven SOC architecture that integrates cloud telemetry, threat intelligence, MITRE ATT&CK Cloud Matrix mappings, and SOAR-based response automation is proposed that demonstrates a 41.2% reduction in time-to-detection, a 37.5% reduction in time-to-response, and 87.9% precision in automated remediation.

Jilika Jithendarnadh, J. Balaraju · 0 citations
Open access Sep 2026

From Reactive Monitoring to Predictive Resilience: AI-Assisted Observability in Financial Platforms

While multi-service architectures are increasing in popularity, the financial industry is being challenged as regulatory requirements continue to grow and customers’ expectations for seamless services have never been higher. Current observability practices, which largely depend on the threshold-based alerting and on re...

Kannan Meiappan · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.