Skip to content

SuperDRI: Graph-Guided Autonomous Troubleshooting for Cloud Databases

Aug 2026 · Proceedings of the VLDB Endowment · 0 citations · 10 references

Abstract

We demonstrate SuperDRI, an autonomous troubleshooting system for large-scale cloud databases such as Microsoft Fabric Data Warehouse. Incident diagnosis requires expert, multi-step reasoning across heterogeneous telemetry and tools, while much of this knowledge remains scattered across documentation and senior engineers' experience. Rather than relying on manually authored playbooks, SuperDRI learns structured troubleshooting workflows directly from historical incident traces through a multi-agent extraction pipeline, incrementally building InvestiGraph, a typed workflow graph of database problems, diagnostic actions, and causal relationships for each domain. At runtime, an agentic graph-traversal mechanism explores InvestiGraph, generates and executes readonly Kusto (KQL) queries against database telemetry, and iteratively reasons over results to reach mitigations. The system further incorporates real-time feedback learning, allowing database engineers to inject domain knowledge into the graph. We illustrate four capabilities: (1) exploring a production-scale InvestiGraph learned from over 1,300 real incidents (10,000+ nodes), (2) observing autonomous troubleshooting with adaptive KQL generation and execution against Fabric DW telemetry, (3) editing the graph and observing how feedback immediately improves subsequent runs, and (4) accessing SuperDRI through multiple MCP-compatible clients. SuperDRI is deployed in production at Microsoft for Azure Data incident management.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.