Skip to content
Conference

Chaos Engineering and Observability Integration for Resilient Cloud Infrastructure

Jul 2026 · 2026 International Conference on Intelligent and Sustainable AI Systems (ICOSAAS) · pp. 1137-1144 · 0 citations · 16 references

Abstract

Cloud computing acts as the core foundation supporting modern digital services, and hosts large-scale applications across five major domains: finance, healthcare, e-commerce, education, and industrial automation. However, its inherent complex, distributed, and dynamic native characteristics make it susceptible to four types of failures: sudden latency spikes, service outages, resource exhaustion, and security vulnerabilities. Traditional monitoring systems, which can only respond to failures after they occur, cannot guarantee the resilience of cloud environments. To address this issue, this paper proposes an integrated framework that combines chaos engineering and observability: chaos engineering injects controlled failures into production-like environments to evaluate a system’s load-bearing capacity, while observability obtains in-depth insights into a system through metrics, logs, distributed tracing, and event analysis. This framework unifies the capabilities of the two types of platforms to realize three core functions: proactive failure detection, automated recovery, and continuous resilience verification. We conducted validation experiments based on Kubernetes-powered containerized microservices, paired with Prometheus, Grafana, Jaeger, and LitmusChaos. Experimental results show that the framework achieves notable improvements across four dimensions: failure detection time, system recovery rate, service availability, and operational reliability. It can help all types of organizations identify hidden vulnerabilities, cut downtime, and strengthen service continuity.

View source

Similar papers

Review Open access Sep 2026

SECURE AND RESILIENT CLOUD INFRASTRUCTURE FOR SAUDI ARABIA: INTEGRATING SIEM, AUTOMATED MONITORING, AND INFRASTRUCTURE ANALYTICS

This review examines how Security Information and Event Management, automated monitoring, artificial intelligence for IT operations (AIOps), and infrastructure analytics can be integrated into a single resilience-oriented operating model for Saudi cloud environments.

M. Shaik · 0 citations

Orchestrating Autonomic Software-Defined Networks for Resilience in Critical Energy Infrastructure

Critical energy infrastructure is becoming increasingly digitalized as the world transitions toward cleaner and more sustainable energy systems. Offshore wind power plants and other distributed energy systems rely on industrial communication networks to connect turbines, substations, control centres, and edge computing...

Agrippina Mwangi · 0 citations
Open access Aug 2026

Predictive fault tolerance and autonomous remediation in distributed cloud infrastructure

A three-layer predictive fault tolerance architecture, telemetry collection and signal engineering, machine learning-driven anomaly detection, and autonomous remediation orchestration are presented, designed for the operational realities of production-critical enterprise cloud infrastructure.

Arjun Danda Sureshbabu · 0 citations
Conference Aug 2026

Integrating Observability, DevSecOps, and Policy-as-Code for Enterprise Cloud Reliability

Modern enterprise cloud infrastructures need to be able to deliver software quickly, remain operationally resilient and follow strict regulatory requirements. Through traditional operations, observability, security and governance are typically treated as separate processes, which leads to longer Mean Time To Repair (MT...

Vishwa Lakhnakiya · 0 citations
Open access Sep 2026

Hybrid Mainframe–Cloud Modernization Framework for Resilient and Scalable Financial Infrastructure

Financial institutions operating complex, mission-critical infrastructure face a persistent modernization dilemma: mainframe systems provide hardware-enforced transactional integrity, deterministic performance, and decades of regulatory validation, while cloud infrastructure offers elastic scalability, managed service...

Narasimha Rao Vanaparthi, M. Dhanekula, Rishabh Srivastava · 0 citations
Open access Sep 2026

Reliability Engineering Framework for Cloud Based Government and Financial Information Systems

The accelerated adoption of cloud computing by government institutions and financial organizations has transformed public service delivery, fiscal management, and financial intermediation. However, the migration of mission-critical information systems to cloud environments introduces complex reliability challenges w...

Upreti Mokshada · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.