CUA-Sandbox: Efficient Environments for Computer-Use Agent Reinforcement Learning
CUA-Sandbox is introduced, which separates private state capsules from shared runtimes through state-scoped execution and transactional lifecycle operations, including resets and branches, while retaining the original software interfaces and task evaluators.
Xin Yan, Zheng-Bo Jiao, Jia-Qi Liu et al.
· 0 citations