Skip to content
Open access

Time-Series-Based Self-Healing Architecture for Incident Management in Infrastructure-as-Code

2026 · International Conference on Software and Data Technologies · pp. 236-245 · 0 citations · 25 references
Computer Science

TL;DR

An architecture that connects runtime monitoring, time-series-based anomaly analysis, remediation planning, and IaC-based execution into a codified and auditable feedback loop is presented, providing proof-of-concept evidence that time-series anomaly signals can be linked to codified IaC remediation within a complete detection-to-remediation loop.

Abstract

: This paper presents a closed-loop self-healing architecture for incident management in Infrastructure-as-Code (IaC) environments, structured as an instantiation of the MAPE–K (Monitor–Analyze–Plan–Execute over Knowledge) pattern. The main contribution is an architecture that connects runtime monitoring, time-series-based anomaly analysis, remediation planning, and IaC-based execution into a codified and auditable feedback loop. We first analyze a catalog of 20 IaC incident-management rules to identify which incident types exhibit temporal behavior and may therefore benefit from time-series-based analysis. We then instantiate the architecture for one controlled SSH-related anomaly scenario, where Moving Average (MA) and ARIMA are used as lightweight statistical detectors and Ansible playbooks are used to trigger a temporary ban/unban remediation action. The results provide proof-of-concept evidence that time-series anomaly signals can be linked to codified IaC remediation within a complete detection-to-remediation loop.

Read PDF

Similar papers

Open access Sep 2026

AI-Driven Self-Healing System for Real-Time Software Incident Management

Modern cloud-native software systems, including banking applications, cloud platforms, enterprise resource planning architectures, and real-time online services, continuously produce large volumes of operational data such as application logs, performance metrics, and system traces. Monitoring these systems using conven...

Vijaya Lakshmi G, A. S · 0 citations
Open access Sep 2026

A rule-driven SOC architecture for real-time threat detection and autonomous incident response

Security Operations Centers (SOCs) are essential for monitoring and responding to cyber threats in cloud-native environments, where infrastructure is dynamic, multi-tenant, and API-driven. Conventional SOCs rely heavily on manual triage and SIEM-based alerting, resulting in delayed detection of cloud-specific attacks s...

Jilika Jithendarnadh, J. Balaraju · 0 citations
Review Open access Aug 2026

From Detection to Verified Action: Operational Readiness for AI-Enabled Cloud Failure Management

The proposed Operational Decision-Readiness and Verification framework represents a deployment claim through analytical capability, operational authority, and assurance maturity and is conceptual and requires prospective and inter-rater validation.

Adepegba Akindayomi Akintade · 0 citations
Conference Sep 2026

A cloud-native log analytics-based operating fault tracing method for power monitoring systems

Reliable automated control in cloud-native power monitoring requires faults to be traced before stale measurements, delayed alarms, or unsafe commands propagate through optical-fiber Ethernet links and service chains. This study addresses the research problem of locating the initiating field, communication, or software...

Peng-Yuan Wang, Yi-Min Wang, Xiao-Shuang Lin · 0 citations
Open access 2026

HAMA: A Hierarchical Adaptive Multi-Agent Architecture for Industrial IoT Predictive Maintenance

Industrial IoT predictive maintenance demands real-time anomaly detection under tight resource and interpretability constraints, while monolithic LLM-based systems remain impractical for on-site deployment. We introduce HAMA (Hierarchical Adaptive Multi-Agent Architecture), in which “adaptive” refers strictly to online...

Rebin Saleh, K. Dinh, Balázs Villányi et al. · 0 citations
Conference Aug 2026

Integrating Observability, DevSecOps, and Policy-as-Code for Enterprise Cloud Reliability

Modern enterprise cloud infrastructures need to be able to deliver software quickly, remain operationally resilient and follow strict regulatory requirements. Through traditional operations, observability, security and governance are typically treated as separate processes, which leads to longer Mean Time To Repair (MT...

Vishwa Lakhnakiya · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.