Skip to content

Industrial Practice of LLM-Based Test Case Carving and Assertion Generation (Experience Paper)

Jul 2026 · arXiv.org · Vol abs/2607.24000 · 0 citations · 36 references
Computer Science

TL;DR

NL2Test is presented, an end-to-end approach and tool that generates executable API regression tests from a natural-language scenario description and a traffic capture recorded while executing the scenario, indicating that traffic-grounded generation with deterministic guardrails can substantially reduce manual effort while improving regression automation in complex microservice environments.

Abstract

Enterprise regression testing for microservice systems is often constrained by incomplete or outdated documentation. In practice, QA engineers frequently rely on real execution traffic to reconstruct business scenarios; however, turning raw traffic into replayable regression tests with stable validation logic remains labor-intensive and error-prone. This paper presents NL2Test, an end-to-end approach and tool that generates executable API regression tests from (i) a natural-language scenario description and (ii) a traffic capture recorded while executing the scenario. NL2Test addresses two coupled tasks: test case carving, which extracts a minimal replayable request sequence and reconstructs data dependencies so that dynamic values are bound from their responses rather than hard-coded; and assertion generation, which produces assertions aligned with business intent while avoiding non-deterministic fields and hallucinated paths. To improve reliability, NL2Test uses LLMs for semantic interpretation and constrained code synthesis, and uses deterministic algorithms for request filtering, dependency confirmation via value consistency, and assertion-path validation. We evaluate NL2Test on 51 industrial regression scenarios extracted from a large consumer-facing Internet company. NL2Test achieves an exact-match rate of 82.4% (42/51), and produces a functionally usable draft in 98.0% (50/51) of scenarios when allowing minor post-edits. In a 9-month production deployment starting in March 2025, NL2Test generated 3,196 test cases with an overall code adoption rate of 85.4%. These results indicate that traffic-grounded generation with deterministic guardrails can substantially reduce manual effort while improving regression automation in complex microservice environments.

View source

Similar papers

Conference Jul 2026

Improving LLM-Based Unit Test Generation Through Root-Cause-Driven Prompt Design

This paper addresses automated unit test generation with large language models (LLMs). LLM-based test generation has not yet attained a quality level sufficient for practical use in industry. Although LLMs often reproduce API syntax faithfully, they frequently disregard semantic usage constraints and execution-environm...

Mizuki Yamada, Masahiko Kato, Juichi Takahashi · 0 citations
Preprint Aug 2026

Framework and Benchmark for Code-Driven Agentic Testing in Web Development

CAT is introduced, a paradigm in which the agent writes Playwright code to drive the browser, gathers feedback, and autonomously explores web applications to uncover bugs, revealing a clear gap between current VLM capabilities and the demands of real-world testing in AI web development.

Bin Hong, Zhen-Chao Zhang, Ji-Yuan He et al. · 0 citations
Preprint Aug 2026

Agent-Based Test Assertion Generation via Diverse Perspective Aggregation

AssertMate is proposed, a novel agent-based assertion generation framework that enhances the quality and reliability of LLM-generated assertions through three key components: actual value construction that identifies assertion targets via static analysis and type-aware heuristics, and multi-perspective expected value p...

Dong Wang, Qiaoyu Han, Lin Yang et al. · 0 citations
Preprint Aug 2026

EduPluginBench: Executable Assurance for AI-Generated Educational Plugins

E EduPluginBench is introduced, an executable benchmark and staged admission method for generated plugins in governed software ecosystems that retains protocols, public-source provenance, raw generations, row-level decisions, audits, analysis code, and reproduction instructions.

Nizam Kadir · 0 citations
Preprint Aug 2026

Requirements-Augmented Generation for Trustworthy Acceptance Testing of LLM-Based Software

This paper introduces Requirements-Augmented Generation (REAG), which interprets user intentions by retrieving relevant software requirements, domain knowledge, and personas via adaptive RAG and self-reasoning to generate context-aware test oracles, and introduces a confidence-calibrated cascade judgment, which quantif...

Fan-Yu Wang, Chetan Arora, Zhen-Ping Xie et al. · 0 citations
Jul 2026

Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents

This paper proposes SafeKeep, an inference-time safeguard that decouples safety judgment from tool execution: it assesses requests using flattened textual tool specifications while retaining the original schema-formatted specifications for execution.

Minghui Pan, Jiayuxuan Yang, Yuan-Yuan Yuan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.