Agentic Orchestration for Self-Healing Data Pipelines: A Policy-Aware Framework for Autonomous Schema Drift Detection and Recovery in Cloud-Native Analytics Environments
Abstract
Cloud-native analytics pipelines are vulnerable to schema drift, implicit data-contract violations, and cascading service-level objective (SLO) failures that conventional monitoring or manual DataOps may detect only after downstream impact. This paper presents AegisFlow, a policy-aware agentic orchestration framework for schema-drift detection and bounded autonomous recovery. Its novelty is system-level integration: seven functional layers and five specialized agents combine schema-aware detection, lineage-informed impact analysis, policy-constrained action selection, recovery/rollback, observability, and outcome feedback in one closed loop rather than introducing a new standalone detector. Evaluation is performed in a controlled Kubernetes-based cloud-native emulation, not a production deployment, using four workloads from three public datasets and 1,200 runs. AegisFlow achieves 96.8% drift-detection accuracy, 3.8 min mean recovery latency, 2.4% false-alarm rate, 81.6% downtime reduction versus manual recovery, and 98.6% SLA compliance with 7.2% orchestration overhead. The results support feasibility under controlled conditions, while production traces, larger deployments, higher event rates, and concurrent-failure scenarios remain necessary for external validation.