Skip to content

Author

Qinghao Zhang

We have 6 of 15 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#natural language process... Preprint Sep 2026

Certified Selective Automation of LLM Agent Evaluation

A task-level bootstrap certificate that is valid in every regime the authors test while matching the naive certificate's coverage, and doubles as a self-training filter that lets a judge enter an unseen domain at in-domain strength with zero target training labels.

Cheng-Guang Gan, Yun-Hao Liang, Qing-Hao Zhang et al. · 0 citations
#natural language process... Preprint Sep 2026

Guides That Cause Actions: An Offline Study of Guide-Action Mutual Reinforcement in Multimodal Web Agents

Web agents are usually evaluated in live environments, where environment state and judge models drift between runs, so the same checkpoint rarely reproduces the same score, making controlled studies of training phenomena impractical. We present WebMRE, an offline benchmark of 541 tasks and 5,293 steps derived from succ...

Cheng-Guang Gan, Yun-Hao Liang, Qing-Hao Zhang et al. · 0 citations
#natural language process... Preprint Sep 2026

Joint Training Is Not Enough: Conditioned Cross-Granularity Training for Multimodal Document Understanding

The Mutual Reinforcement Effect is tested in multimodal document understanding on three corpora, two of receipts and one of scanned business forms, comparing single-task, joint and conditioned training, which puts one granularity's gold output in the other's prompt during training only.

Cheng-Guang Gan, Yun-Hao Liang, Han-Jun Wei et al. · 0 citations
Jul 2026

A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism

This work asks whether it adds skill to a small language and vision-language model web agent at the 4B to 8B scale, or whether it mostly reshapes behavior the supervised model already has, and explains the failure of GRPO.

Cheng-Guang Gan, Zhi-Xi Cai, Yun-Hao Liang et al. · 0 citations
Jul 2026

MAG: A Web-Agent Benchmark and Harness for Multimodal Action and Guide Generation

MAG is introduced, the first benchmark that unifies task execution and guide writing into a single Multimodal Action and Guide task, with two grounding schemes over screenshots: Set-of-Mark element selection and raw pixel coordinates.

Chengguang Gan, Hanjun Wei, Yun-Hao Liang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.