Skip to content
Preprint

Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds

Aug 2026 · 0 citations · 63 references
Computer Science

TL;DR

It is found that visual presentation is a central factor that determines whether VLM benchmarks measure grounded perception, downstream reasoning, or a mixture of both, and gains are closely tied to reductions in grounding-related errors, while rule reasoning remains comparatively challenging.

Abstract

Vision-language models (VLMs) have advanced rapidly in multimodal reasoning, yet recent work shows that their failures often reflect an interaction between visual grounding and downstream reasoning. What remains less clear is how the visual presentation of a task shapes model performance and failure modes when the underlying reasoning problem is unchanged. We study this question in SPaRC, a benchmark for grid-based visual spatial planning, by introducing lightweight input-side scaffolds that preserve the visual modality while making spatial structure more accessible. Across multiple VLMs, these scaffolds improve task accuracy over the original visual setting by up to 34.0 percentage points and further complement GRPO-based training, yielding up to 4.6 additional accuracy points compared with near-zero gains on the original visual input. Analyses on both end-to-end task solving and object detection show that these gains are closely tied to reductions in grounding-related errors, while rule reasoning remains comparatively challenging. We find that visual presentation is a central factor that determines whether VLM benchmarks measure grounded perception, downstream reasoning, or a mixture of both.

View source

Similar papers

Jul 2026

SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them

SpatialCLI is proposed, a framework that teaches VLMs to reason with spatial tools and progressively internalize the specialist perceptual capabilities they provide and introduces SpatialCLI-Bench, a 516-example benchmark for compositional perception across localization, segmentation, depth, and pose.

Yang Zhou, Zixuan Huang, Sunzhu Li et al. · 0 citations
Jul 2026

See2Think: Do Multimodal Models Really Use Intermediate Visual States?

Evaluating representative proprietary and open-source multimodal models, it is found that visual reasoning is strongly model- and environment-dependent, with no single setting consistently dominating across tasks.

Siyu Yan, Zhuoran Yan, Haiying Xu et al. · 0 citations
#artificial intelligence Preprint Aug 2026

LOCI: A Locator-Critic with Refinement Loop

Locator-Critic (LOCI) is proposed, a training-free framework that decouples visual search from evidence verification and improves accuracy for both open-weight models like Qwen3-VL and proprietary models like Gemini 2.5 Pro.

Walid Bousselham, Mathilde Caron, Arsha Nagrani et al. · 0 citations
Preprint Aug 2026

Think with Structured Grounding: Perceptual Reinforcement Learning for Chart and Visual-Tabular Understanding

Think with Structured Grounding (TwSG), a novel fine-grained image perception framework designed to internalize complex images's tool-use capabilities within the model, is proposed, endowing models with native fine-grained region description and flexible reasoning capabilities.

Chang-Jiang Jiang, Qiannian Zhao, Lei Xin et al. · 2 citations
Preprint Aug 2026

When Does Visual Generation Help Visual Understanding in Unified Multimodal Models?

VGAU-Diag is introduced, a fine-grained evaluation framework for vision generation-assisted understanding that stratifies samples by difficulty, enables unified evaluation of multiple reasoning paradigms, and uses Oracle-Ass Reference Protocols.

Yu-Bo Zhu, Zhe-Han Kan, Jing-Yi Yang et al. · 1 citation
Jul 2026

Spatial-IQ: Deconstructing Spatial Intelligence via Hierarchical Capability Tests

It is demonstrated that training models with chain-of-thought supervision over the authors' hierarchical sub-tasks, combined with reinforcement learning with verifiable rewards, significantly improves both spatial consistency across sub-tasks and target-task accuracy, supporting the value of the proposed decomposition...

Patrick Rim, Tom Long, Ekta Prashnani et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.