VLM-in-Sandbox: Visual Workspaces for Agentic Visual Reasoning
Across seven benchmarks and four base VLMs, VLM-in-Sandbox achieves the highest sample-weighted average accuracy among Vanilla VLM, Append-only Sandbox, and the proposed method, and is identified as a central abstraction for sandboxed VLM agents.
He-Xiong Yang, Mingrui Chen, Jie Cao et al.
· 0 citations