Skip to content
Preprint

From Bug Reports to Browser-Executable Procedures: An LLM-Driven Agent for Web GUI Bug Reproduction

Aug 2026 · 0 citations · 53 references
Computer Science

TL;DR

The results show that explicit context reconstruction and state-aware browser execution effectively support report-derived browser reproduction, while historical replay shows that successful procedures often expose the original bug-present behavior on restored buggy versions.

Abstract

Reproducing web GUI bugs from natural-language bug reports is critical for software maintenance, but remains difficult because reports often lack prerequisites such as dependencies and input files. Existing bug reproduction techniques mainly target code units or mobile applications and lack end-to-end visual execution and validation for web GUIs. We present ReBug, a context-aware agent system that reconstructs, executes, and validates browser-level reproduction procedures from web GUI bug reports by driving a real browser. ReBug separates reproduction into two stages. In the preparation stage, ReBug reconstructs missing prerequisites from the report and available artifacts, and it produces a high-level reproduction plan. In the execution stage, it performs tool-mediated interactions in the browser, maintains structured summaries of page state and action history, and validates the final state against expectations derived from the report. We evaluate ReBug on 667 real-world bug reports from four open-source web applications. On controlled current deployments, ReBug outperforms both baselines, achieving an average RSR of 49.96%, a mean task completion rate of 74.96%, and a mean action execution success rate of 86.54%. Our results show that explicit context reconstruction and state-aware browser execution effectively support report-derived browser reproduction, while historical replay shows that successful procedures often expose the original bug-present behavior on restored buggy versions.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

WatchPoint: Executable User Feedback for Real-World Agentic Web Development

When a professional web developer's code fails a test, they do not simply re-read the stack trace. They open the application in a browser, click buttons, inspect computed styles, and run diagnostic commands to understand what went wrong. Existing feedback mechanisms for coding agents rely on screenshots, LLM-as-a-judge...

Guanqun Yang, Wei Yang, Xueqing Liu · 0 citations
Preprint Aug 2026

Framework and Benchmark for Code-Driven Agentic Testing in Web Development

CAT is introduced, a paradigm in which the agent writes Playwright code to drive the browser, gathers feedback, and autonomously explores web applications to uncover bugs, revealing a clear gap between current VLM capabilities and the demands of real-world testing in AI web development.

Bin Hong, Zhen-Chao Zhang, Ji-Yuan He et al. · 0 citations
Open access Aug 2026

From Logging Configuration to Code Execution: A Systematization of Log4j 2 File-Write Primitives in HTTP-Exposed JMX

Java middleware may expose Java Management Extensions (JMX) through Jolokia’s Hypertext Transfer Protocol (HTTP) bridge. In affected ActiveMQ deployments, reachable Log4j 2 configuration managed beans (MBeans) become write capabilities and, with compatible triggers, enable remote code execution (RCE). We ask: in a spec...

A. Caciulescu, Matei Badanoiu, R. Rughinis et al. · 0 citations
#natural language process... Preprint Aug 2026

WebWorld: The Browser as a World Model for Self-Improving Web Code

WebWorld is presented, the interface that lets a VLM prior interact with this browser-as-world-model autonomously and decides which interactions become supervision and reaches the level of strong frontier systems such as Kimi-K2.6 and GPT-5.4 on interactive HTML generation.

Jia-Jun Wu, Jian Yang, Ya-Xin Du et al. · 0 citations
#human-computer interacti... Book Open access Aug 2026

FlowCheck: Helping End-Users Specify and Verify Intent in Vibe-Coded Web Apps

FlowCheck, a constraint language to specify user-visible information flows directly through the application interface, where constraints can also be displayed and inspected without reading code, and are structured enough for reliable LLM generation.

Reya Vir, Lydia B. Chilton, Zhuo Zhang et al. · 2 citations
Review Sep 2026

Test-Driven Approaches to Software Engineering with Large Language Models: A Survey of Phases, Tasks, and Agent Skills

A structured scoping survey organized around the question of what decision a test changes is presented, and distinguishes the Red--Green--Refactor cycle from test-conditioned generation, execution-guided refinement, test-mediated analysis, and evaluation-only testing.

Yun-Hao Liang, Cheng-Guang Gan, Rui-Xuan Ying et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.