Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds
It is found that visual presentation is a central factor that determines whether VLM benchmarks measure grounded perception, downstream reasoning, or a mixture of both, and gains are closely tied to reductions in grounding-related errors, while rule reasoning remains comparatively challenging.