Jul 2026· Anais do LIII Seminário Integrado de Software e Hardware (SEMISH 2026)· pp. 191-202· 0 citations· 18 references
TL;DR
Empirical evidence is provided that MCP-orchestrated LLM agents can support root cause analysis in cloud-native environments and offer practical guidance for model selection in AIOps/SRE workflows.
Abstract
Context: The growing complexity of cloud microservices imposes significant challenges for Site Reliability Engineering (SRE), contributing to delayed incident resolution and increased operational effort. Objective: This study evaluated the effectiveness of autonomous agents based on Large Language Models (LLMs), orchestrated via the Model Context Protocol (MCP), for root cause analysis in a cloud-native setting. Method: We conducted a controlled Randomized Complete Block Design (RCBD) experiment in Kubernetes with automated fault injection, covering three distinct failure scenarios and multiple LLM configurations across 360 executions. Results: A high-performing configuration (Gemini 2.5 Flash at low temperature) achieved a 71.1% root-cause identification success rate, substantially above a random-chance baseline (≈ 0.91%). Smaller models exhibited higher token and step volatility (CV = 2.17) and more repeated tool-call cycles, challenging the assumption that lower-parameter models are inherently more cost-effective for SRE workflows. Conclusion: The results provide empirical evidence that MCP-orchestrated LLM agents can support root cause analysis in cloud-native environments and offer practical guidance for model selection in AIOps/SRE workflows.
The results suggest that process-level orchestration combined with automated monitoring components can serve as a practical alternative for hosting small Express.js services.
F. A. Desyatirikov· PROGRAMMNAYA INGENERIA· 0 citations
Examination of tool-augmented Large Language Model systems for supporting Root Cause Analysis of nightly test failures at Westermo Network Technologies AB finds the single agent system generated reports faster and at lower cost, making it the more practical baseline in this context.
Eric Jansson, P. Strandberg, Thomas Sörensen et al.· 0 citations
The swift advancement of cloud-native computing has greatly heightened the complexity involved in managing application reliability, infrastructure governance, and continuous software delivery within highly distributed environments. While recent developments in Infrastructure as Code (IaC), observability, AIOps, and AI-...
M. Dhanekula· International Journal for Sc...· 0 citations
Multi-agent orchestration frameworks are moving from demos to production, yet benchmarks typically report task accuracy without diagnosing why a pipeline failed, where a cascade began, or which routing decision caused the breakdown. OrchestraBench evaluates failure, recovery, and decomposition through a controlled, see...
Yidian Chen, Ying Gu, Natan Vidra et al.· 0 citations
Cloud–edge computing enables scalable and resilient deployment of microservice-based applications, however achieving resource efficiency while ensuring stringent Quality of Service (QoS) remains challenging. The strong interdependencies among microservices and non-linear latency effects near resource saturation render...
Lazaros Liatsas, Godfrey M. Kibalya, Angelos Antonopoulos· IEEE Transactions on Network...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.