Preprint
Jul 2026
MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents
Evaluating 12 state-of-the-art models, from 4B open-weight to frontier proprietary systems, shows that current models still lack robust visual tool-calling capability: even the best model achieves below 50% success rate, suggesting fundamentally different research directions for improving models at different capability levels.
Kaixin Ma, Di Feng, Alexander Metz et al.
· 0 citations