Skip to content
Preprint

WebGrader: Training LLMs for Web Development with Self-Evolving Programmatic Grader

Aug 2026 · 0 citations · 26 references
Computer Science

TL;DR

WebGrader is proposed, a self-evolving programmatic grader that autonomously derives the required interaction flows from each website request, represents each flow as an executable Flow Contract, and uses its execution outcome as an RL reward.

Abstract

Large language models increasingly generate complete websites from natural-language descriptions, and reinforcement learning has become a central approach to closing their remaining functional gap. This training regime is bottlenecked by reward design. Hand-authored browser scripts are executable yet costly to write for open-ended requirements, while VLM and GUI-agent graders scale but may issue verdicts before observing the decisive state. We propose WebGrader, a self-evolving programmatic grader that autonomously derives the required interaction flows from each website request, represents each flow as an executable Flow Contract, and uses its execution outcome as an RL reward. WebGrader materializes the generated project in a live browser, grounds target actions against the source code and live DOM, and collects visual, DOM, response, and persistent-state evidence along the same browser trajectory. A residual-driven offline loop then discovers reusable verifier skills, screens them on disjoint validation pages, and freezes the promoted skill graph before policy training. By separating test planning, action grounding, evidence collection, and semantic judgment, WebGrader issues a Pass verdict only after observing the requested transition. On WebGen-Bench, WebGrader trains an 8B policy to a 52.01% functional success rate, outperforming a matched appearance-plus-script reward by 7.88 points and surpassing o4-mini and DeepSeek-v4-flash. On WG-core-250, the policy reaches a Full Score of 44.953 and surpasses Qwen3-Coder-480B.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

WatchPoint: Executable User Feedback for Real-World Agentic Web Development

When a professional web developer's code fails a test, they do not simply re-read the stack trace. They open the application in a browser, click buttons, inspect computed styles, and run diagnostic commands to understand what went wrong. Existing feedback mechanisms for coding agents rely on screenshots, LLM-as-a-judge...

Guanqun Yang, Wei Yang, Xueqing Liu · 0 citations
#natural language process... Preprint Aug 2026

WebWorld: The Browser as a World Model for Self-Improving Web Code

WebWorld is presented, the interface that lets a VLM prior interact with this browser-as-world-model autonomously and decides which interactions become supervision and reaches the level of strong frontier systems such as Kimi-K2.6 and GPT-5.4 on interactive HTML generation.

Jia-Jun Wu, Jian Yang, Ya-Xin Du et al. · 0 citations
Preprint Aug 2026

Framework and Benchmark for Code-Driven Agentic Testing in Web Development

CAT is introduced, a paradigm in which the agent writes Playwright code to drive the browser, gathers feedback, and autonomously explores web applications to uncover bugs, revealing a clear gap between current VLM capabilities and the demands of real-world testing in AI web development.

Bin Hong, Zhen-Chao Zhang, Ji-Yuan He et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Rubric-to-Code Credit Assignment for Reinforcement Learning

A reinforcement learning framework that converts rubric-level functional feedback into localized optimization signals over generated code and aligns evaluator-generated textual attributions with responsible code spans and generated tokens is proposed.

Rui Jin, Jikai Chen, Yihang Chen et al. · 0 citations
Review Aug 2026

LiveEvalBench: Toward Open-World Evaluation for Web Generation

Experiments across diverse real-world web-generation scenarios show that LiveEvalBench aligns closely with human expert judgment and provides fine-grained insights into frontier models' web generation capabilities.

Yiyao Wang, Zhen Wen, Ying Tang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.