Skip to content
Preprint

HxAgent: Iterative Agent Planning for End-to-End Web Application Testing

Aug 2026 · 0 citations · 49 references
Computer Science

TL;DR

HxAgent is introduced, an iterative LLM-based planning agent with a proactive correction strategy that achieves 97.4% Exact-Match accuracy on MiniWoB++, comparable to the best baselines without human demonstrations and surpassing the recent WALT by 10.5%.

Abstract

In automated web testing, generating test cases and performing testing using functionality descriptions in natural-language is crucial for improving efficacy. These tasks require such a testing agent to carry out tasks on the target application and generating tests autonomously. We introduce HxAgent, an iterative LLM-based planning agent with a proactive correction strategy. After each step, HxAgent reassesses the web state to determine the next action using (1) current observations, (2) short-term memory of past actions, and (3) long-term experience extracted from past (in)correct sequences of actions. HxAgent achieves 97.4% Exact-Match accuracy on MiniWoB++, comparable to the best baselines without human demonstrations and surpassing the recent WALT by 10.5%. On a dataset of 350 web tasks, it attains 83.8% Exact-Match and 91.8% Prefix-Match, exceeding WALT by 13.4%. On OnlineMind2Web, it further improves over WALT by 4.6%.

View source

Similar papers

Conference Open access Sep 2026

PENTESTLLMAGENT: A Task Dependency Graph Planning-Based Multi-Agent Framework for Automated Penetration Testing

PentestLLMAgent is proposed, which integrates a Task Dependency Graph (TDG) for dynamic planning and backtracking; a Hierarchical Multi-Agent Architecture (HMA) with function-calling-based tool invocation, output filtering, and semantic compression, and Executable Knowledge-Guided Command Generation (EKG-CG) for retrie...

Shuo Sheng, Jixin Zhang, Jia Yang et al. · 0 citations
Preprint Aug 2026

An Actionable Diagnosis of Multilingual, Multi-Agent Planning Failures

To test whether the taxonomy supports mitigation, TART, Taxonomy-Guided Actionable Representation, is introduced that makes the taxonomy's key aspects explicit to the planner and downstream sub-agents and consistently improves performance.

Vikas Pahuja, J. Brokman, O. Hofman et al. · 0 citations
Preprint Aug 2026

TDD-Agent: Test-Driven Reasoning for Code Generation

TDD-Agent is introduced, which operationalizes the test-driven development paradigm for code generation and improves not only code correctness but also the effectiveness of the generated tests, yielding higher pass rates, coverage, and mutation scores, suggesting that tests can serve as evolving reasoning artifacts rat...

Hong Yu, Ke-Fan Li, Jia-Kun Li et al. · 2 citations
#natural language process... Preprint Aug 2026

A^2Agent: Action-Aware Reinforcement Learning for Repository-Level Code Localization Agents

An action-aware reinforcement learning method that combines a per-turn reward sequence rewarding both the discovery and commitment of gold code regions with an action-level advantage estimation scheme that isolates each action's credit by grouping turns sharing the same exploration context is proposed.

Doyeon Kim, Suyoung Bae, Yu-Min Lee et al. · 1 citation
#natural language process... Preprint Sep 2026

MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding

Recently, the rapid development of large language models (LLMs) has reshaped software engineering by enabling autonomous code agents that plan, execute, and utilize external tools iteratively to tackle complex tasks. Beyond achieving functional correctness, these agents must faithfully follow process instructions and c...

Bo-Si Wen, Cunxiang Wang, Jia-Yi Gui et al. · 0 citations
#artificial intelligence Preprint Sep 2026

TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces

An agent can complete a task while exhibiting undesirable behavior during execution. Developers need tests for the specific behaviors encountered in deployment, beyond fixed benchmark suites. We present TraceDance, an agent system that constructs targeted benchmarks from deployment traces for user-specified undesirable...

De-Hai Min, Dao-An Zhang, Yiming Zeng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.