It is found that visual presentation is a central factor that determines whether VLM benchmarks measure grounded perception, downstream reasoning, or a mixture of both, and gains are closely tied to reductions in grounding-related errors, while rule reasoning remains comparatively challenging.
Abstract
Vision-language models (VLMs) have advanced rapidly in multimodal reasoning, yet recent work shows that their failures often reflect an interaction between visual grounding and downstream reasoning. What remains less clear is how the visual presentation of a task shapes model performance and failure modes when the underlying reasoning problem is unchanged. We study this question in SPaRC, a benchmark for grid-based visual spatial planning, by introducing lightweight input-side scaffolds that preserve the visual modality while making spatial structure more accessible. Across multiple VLMs, these scaffolds improve task accuracy over the original visual setting by up to 34.0 percentage points and further complement GRPO-based training, yielding up to 4.6 additional accuracy points compared with near-zero gains on the original visual input. Analyses on both end-to-end task solving and object detection show that these gains are closely tied to reductions in grounding-related errors, while rule reasoning remains comparatively challenging. We find that visual presentation is a central factor that determines whether VLM benchmarks measure grounded perception, downstream reasoning, or a mixture of both.
SpatialCLI is proposed, a framework that teaches VLMs to reason with spatial tools and progressively internalize the specialist perceptual capabilities they provide and introduces SpatialCLI-Bench, a 516-example benchmark for compositional perception across localization, segmentation, depth, and pose.
Yang Zhou, Zixuan Huang, Sunzhu Li et al.· arXiv.org· 0 citations
Evaluating representative proprietary and open-source multimodal models, it is found that visual reasoning is strongly model- and environment-dependent, with no single setting consistently dominating across tasks.
Locator-Critic (LOCI) is proposed, a training-free framework that decouples visual search from evidence verification and improves accuracy for both open-weight models like Qwen3-VL and proprietary models like Gemini 2.5 Pro.
Walid Bousselham, Mathilde Caron, Arsha Nagrani et al.· 0 citations
Think with Structured Grounding (TwSG), a novel fine-grained image perception framework designed to internalize complex images's tool-use capabilities within the model, is proposed, endowing models with native fine-grained region description and flexible reasoning capabilities.
Chang-Jiang Jiang, Qiannian Zhao, Lei Xin et al.· 2 citations
VGAU-Diag is introduced, a fine-grained evaluation framework for vision generation-assisted understanding that stratifies samples by difficulty, enables unified evaluation of multiple reasoning paradigms, and uses Oracle-Ass Reference Protocols.
Yu-Bo Zhu, Zhe-Han Kan, Jing-Yi Yang et al.· 1 citation
It is demonstrated that training models with chain-of-thought supervision over the authors' hierarchical sub-tasks, combined with reinforcement learning with verifiable rewards, significantly improves both spatial consistency across sub-tasks and target-task accuracy, supporting the value of the proposed decomposition...
Patrick Rim, Tom Long, Ekta Prashnani et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.