This work introduces ScaleCUA, a unified framework that scales online RL for CUAs via verifiable task synthesis and efficient training, and designs VeriGen, an end-to-end framework for generating verifiable RL tasks through iterative docker interactions and a multi-agent feedback loop.
Abstract
Computer use agents (CUAs) are emerging as a powerful interface for automating complex digital workflows through visual perception and GUI execution. Online reinforcement learning with verifiable rewards (RLVR) has emerged as a key direction for scaling their capabilities. However, this paradigm is bottlenecked by verifiable data scarcity and online RL inefficiency. To break these barriers, we introduce ScaleCUA, a unified framework that scales online RL for CUAs via verifiable task synthesis and efficient training. At the data level, we design VeriGen, an end-to-end framework for generating verifiable RL tasks through iterative docker interactions and a multi-agent feedback loop. Scaled to 100+ concurrent agent workers via a shared docker interaction probe, this pipeline produces 24K+ verifiable tasks and nearly 3K high-quality RL tasks. To maximize sample efficiency, we propose Frontier Sampling, which tracks per-task capability and allocates rollouts to the current learning frontier. On the training side, we further design Visual Context Segmentation, a sliding window over recent visual context that balances rollout and training-engine pressure, yielding a 2.83x training speedup over step-wise decomposition. Together, ScaleCUA achieves 68.7% on OSWorld and 54.0% on ScienceBoard, establishing new state-of-the-art performance among open-source computer use agents. Code, models, and datasets are available at https://github.com/THUDM/SCALE-CUA.
This report presents HyMobileAgent, a mobile GUI agent built on Hy3.0-VL-A3B, a vision-native foundation model featuring native any-resolution input, an A3B-scale deployment budget, and a 32K context window to model extended interaction histories.
H. Team, Hua-Wen Shen, Zheng-Yang Tang et al.· arXiv.org· 2 citations· ⚡1
Qwen-CUA is introduced, a native computer-use agent with a 397B-A17B Qwen mixture-of-experts backbone that outperforms Qwen3.7 and remains competitive with leading proprietary systems, and scalable verifiable interaction and hybrid tool use as key directions.
Dunjie Lu, Shuai Bai, Tianyi Bai et al.· 2 citations
GSAR (Goal-State-Anchor Reward), a RL reward framework that supports scalable task generation and delivers reliable reward signals for stable and efficient policy optimization, is introduced.
Long Zhang, Yuhan Chen, Chaoran Zhang et al.· 0 citations
This work proposes Reinforcement Learning with Human-Engine Verification (RLHEV), a post-training paradigm that combines dense engine signals with implicit human acceptance feedback from the development process to support RL post-training.
P. Zhou, Hesong Wang, Zhengfeiyang Zhang et al.· 0 citations
Experimental results demonstrate that the ToolVerse framework significantly strengthens LLMs'capabilities in long-horizon tool use, achieving a marked performance boost and showcasing robust reasoning within dynamic environments.
Shuaiyu Zhou, Fengpeng Yue, Zengjie Hu et al.· arXiv.org· 2 citations
Experiments on GUI agent benchmarks show that LACL-GUI provides more effective learning signals and consistently improves agent performance over prior methods, highlighting the value of trajectory-level supervision in contrastive RLVR.
Chengyang Gu, Le Zhang, Jing-Bo Zhou et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.