This work uses the Qwen3-VL-4B (Bai et al., 2025), a modern VLM, to decode how visual information is encoded across depths and proposes a synthetic dataset of simple geometric shapes for controlled analysis, along with queries crafted to precisely test language cues.
Abstract
Vision-Language Models (VLMs) achieve strong performance in visual reasoning tasks, but it remains unclear whether they understand visual relations, or simply employ shortcuts such as language cues or priors. To investigate this, we use the Qwen3-VL-4B (Bai et al., 2025), a modern VLM, to decode how visual information is encoded across depths. For this, we propose a synthetic dataset of simple geometric shapes for controlled analysis, along with queries crafted to precisely test language cues. Furthermore, the dataset is modified to test causal reliance on visual evidence. Our results show that current VLMs combine genuine visual reasoning with shortcut strategies primarily rooted in language cues.
This work observes that highly causal vision tokens often lie outside the target region, and extends the analysis to larger vision-language models, suggesting that strong multimodal performance does not necessarily imply spatially localized causal representations.
S. NarenKumar, T. Bhatt, Mayank Singh· 0 citations
For the coarse attributes the authors study, MLLMs encode the visual evidence but cannot reliably control their reliance on it, indicating that for the coarse attributes they study, MLLMs cannot reliably control their reliance on it.
Jiaang Li, Chengzu Li, Zhaochong An et al.· arXiv.org· 0 citations
It is found that visual presentation is a central factor that determines whether VLM benchmarks measure grounded perception, downstream reasoning, or a mixture of both, and gains are closely tied to reductions in grounding-related errors, while rule reasoning remains comparatively challenging.
Lars Benedikt Kaesberg, Tianyu Yang, F. Wunderlich et al.· 0 citations
SpatialCLI is proposed, a framework that teaches VLMs to reason with spatial tools and progressively internalize the specialist perceptual capabilities they provide and introduces SpatialCLI-Bench, a 516-example benchmark for compositional perception across localization, segmentation, depth, and pose.
Yang Zhou, Zixuan Huang, Sunzhu Li et al.· arXiv.org· 0 citations
Plane geometry remains a significant challenge in AI, requiring the integration of visual perception and mathematical reasoning. While Large Multimodal Models (LMMs) naturally handle visuo-linguistic inputs, they are often computationally intensive and opaque. We demonstrate that a pure Large Language Model (LLM), when...
Wei-Chen Dai, Rafael Cabral, Ziyi Shou et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.