TSGen, an automated pipeline for generating high-quality, structured TSGs from historical incident reports using large language models (LLMs), consists of filtering and classifying incident data into diagnostically relevant categories and distilling core incidents to ensure diversity and generalizability.
Abstract
Maintaining up-to-date troubleshooting guides (TSGs) is critical for the reliability of cloud systems, yet manual maintenance often leads to inefficiencies and outdated documentation. This paper proposes TSGen, an automated pipeline for generating high-quality, structured TSGs from historical incident reports using large language models (LLMs). Our approach consists of three stages: (1) filtering and classifying incident data into diagnostically relevant categories, (2) distilling core incidents to ensure diversity and generalizability, and (3) organizing the distilled knowledge into a directed acyclic graph (DAG) that captures root causes and resolutions in a structured manner. By leveraging real-world incident discussions, TSGen produces dynamic and reusable guides tailored for live troubleshooting. Experiments on real-world incidents from Microsoft demonstrate that TSGen achieves 54.8% incident coverage and approximately 3× higher retrieval accuracy compared to baselines. Furthermore, the system supports iterative updates, allowing guides to evolve alongside dynamic cloud environments. Human evaluation shows that on-call engineers rate these generated TSGs significantly higher than human-crafted ones.
This work introduces \textsc{JailbreakSkill}, a skill-centric framework for scaling automated red-teaming through reusable and continuously evolving attack capabilities, which packages existing attack strategies into modular, agent-ready skills that can be directly reused and adaptively selected across tasks and target...
Xiaoyu Wen, Jiajia Li, Zhida He et al.· 2 citations
Container images are fundamental to cloud deployment, with their build instructions (e.g., Dockerfiles) critically impacting the efficiency and stability of cloud service. Manually authoring these instructions is error-prone, while Large Language Models (LLMs) lack the domain knowledge to generate both correct and opti...
Kun Wang, Yao Wu, Hao Fan et al.· Fall Joint Computer Conferen...· 0 citations
We demonstrate SuperDRI, an autonomous troubleshooting system for large-scale cloud databases such as Microsoft Fabric Data Warehouse. Incident diagnosis requires expert, multi-step reasoning across heterogeneous telemetry and tools, while much of this knowledge remains scattered across documentation and senior enginee...
Yi-Wen Zhu, Joyce Cahoon, Qiushi Bai et al.· Proceedings of the VLDB Endo...· 0 citations
E EduPluginBench is introduced, an executable benchmark and staged admission method for generated plugins in governed software ecosystems that retains protocols, public-source provenance, raw generations, row-level decisions, audits, analysis code, and reproduction instructions.
Distributed tracing is essential for observing microservices but incurs prohibitive storage costs. Existing solutions primarily rely on trace-level sampling, which indiscriminately discards entire traces based on probability or tail latency. This coarse-grained approach often loses critical off-the-path anomalies and f...
Yulun Wu, Guangba Yu, Zhihan Jiang et al.· ACM Transactions on Software...· 0 citations
The approach formulates repair as a supervised sequence generation task and uses format tags, oracle validation, and boundary-localized repair to generate valid outputs while preserving content, showing strongest content preservation when repairs are successful.
Ovi Paul, Tom J King, Ali Shokri· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.