Examination of tool-augmented Large Language Model systems for supporting Root Cause Analysis of nightly test failures at Westermo Network Technologies AB finds the single agent system generated reports faster and at lower cost, making it the more practical baseline in this context.
Abstract
This study examines tool-augmented Large Language Model (LLM) systems for supporting Root Cause Analysis (RCA) of nightly test failures at Westermo Network Technologies AB. Nightly test executions produce heterogeneous test data and logs that practitioners currently inspect manually across multiple sources. We implemented an RCA workflow in single-agent and orchestrated multi-agent configurations, both with access to test metadata and logs. An exploratory industrial case study used two real failure scenarios. Six practitioners evaluated the scenario reports through a survey and focus group, and operational measurements were collected from 120 repeated executions. The evaluation covered practitioner-perceived correctness, reasoning quality, fix realism, clarity, usefulness, and trust, as well as cost, duration, and consistency. Neither configuration showed a consistent practitioner-perceived quality advantage across the two scenarios. The single agent system generated reports faster and at lower cost, making it the more practical baseline in this context. The potential benefits of agent architectures require further evaluation in more complex scenarios.
DDBench is introduced, a code-repair benchmark of 60 historical bugs mined from 13 open-source distributed systems, partitioned into three difficulty tiers, isolating the effect of debugging context from model capability.
Yi-Bo Yan, Huijuan Wang, Jun-Zhou He et al.· 0 citations
As large language models evolve from question-answering systems into general-purpose agents, evaluation must move beyond static answer correctness to assess multimodal perception, multi-step execution, tool use, and artifact delivery. However, existing benchmarks are often tied to specific task types, execution environ...
Yu Liu, Zhi-Lin Liu, Zhi-Wei Yang et al.· 0 citations
Tool-using agents can fail twice: a required tool can fail, and the agent can then report success without the evidence needed to justify it. Existing benchmarks often entangle this reporting failure with tool selection, recovery, and environment dynamics. We introduce Failure-Transparent Agents (FTA), a controlled benc...
Jun-Ru Zhu, Shi-Ming Xie, Ai-Me-Lu-Fan Chen et al.· 0 citations
A large-scale empirical study of quality assurance (QA) practices in 157 open-source LLM-based agent projects with at least 100 GitHub stars highlights the need to move beyond feature-level testing toward systematic end-to-end validation that ensures agent workflows remain within intended boundaries when interacting wi...
Wu-Yang Dai, Moses Openja, Jiho Shin et al.· 0 citations
The proliferation of large language model (LLM)-based autonomous agents has created a new class of distributed system: the multi-agent LLM network. While significant research focuses on the intelligence of individual agents, comparatively little work addresses the software architectural concerns that govern how fleets...
Ketankumar Savajiyani· 2026 International Conferenc...· 0 citations
Tool-using language-model agents select and execute third-party artifacts. Different implementations can return the requested output while producing hidden execution effects that task-, attack-, or choice-based evaluations may miss. We study functional counterfeits: implementations that match benign alternatives on the...
XiaoYu Xu, Zi Liang, Min-Xin Du et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.