Jun 2026
OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
These results show that current agents are still far from professional-level computer use: rather than stumbling on basic GUI control or coding, they lose track of constraints, miss information that arrives mid-task, guess rather than ask the user, and skip verification, struggling most when a task hinges on hidden state they must recover.
Mengqi Yuan, Zilong Zhou, Xinzhuang Xiong et al.
· arXiv.org · 11 citations
· ⚡5