Skip to content

When Depth Is Better Told Than Shown: Depth-Ordinal Prompting for Vision-Language Spatial Reasoning

Jul 2026 · arXiv.org · Vol abs/2607.11173 · 0 citations · 35 references
Computer Science

TL;DR

This work proposes Depth-Ordinal Prompting (DOP), a training-free method that converts monocular depth into a single question-targeted ordinal text cue at the queried objects, without adding a depth image, training a module, injecting features, or using labels.

Abstract

Vision-language models (VLMs) are expected to reason about physical space -- which object is closer, what lies behind what, and how objects are arranged in 3D -- yet they still struggle with such spatial judgments. A natural remedy is to show the model a depth map, but we find that this can make performance worse. We show that depth is not absent: it reaches the language model, but becomes difficult to access for downstream reasoning, while rendered pseudo-depth maps act as noisy auxiliary images that frozen VLMs cannot easily regulate. We propose Depth-Ordinal Prompting (DOP), a training-free method that converts monocular depth into a single question-targeted ordinal text cue at the queried objects, without adding a depth image, training a module, injecting features, or using labels. Our key finding is form dependence: the same depth signal can hurt when shown as an image but help when told as text.Across benchmarks, models, and depth estimators, DOP improves spatial reasoning when pseudo-depth provides reliable object-level ordering and remains largely neutral in strong original-image regimes. It is also competitive with the strongest training-free depth-prompting alternative while being simpler and more targeted.

View source

Similar papers

Jul 2026

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

ProVisE (Protocolized Visual Evaluation), a benchmark-agnostic framework that elicits protocol-constrained visual answers from image-generation models and parses them into structured predictions compatible with original metrics, is proposed and revealed, revealing complementary strengths of pixel-space expression and t...

Xu Wang, Kaixiang Yao, Miao Pan et al. · 1 citation
Jul 2026

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

For the coarse attributes the authors study, MLLMs encode the visual evidence but cannot reliably control their reliance on it, indicating that for the coarse attributes they study, MLLMs cannot reliably control their reliance on it.

Jiaang Li, Chengzu Li, Zhaochong An et al. · 0 citations
Preprint Aug 2026

Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds

It is found that visual presentation is a central factor that determines whether VLM benchmarks measure grounded perception, downstream reasoning, or a mixture of both, and gains are closely tied to reductions in grounding-related errors, while rule reasoning remains comparatively challenging.

Lars Benedikt Kaesberg, Tianyu Yang, F. Wunderlich et al. · 0 citations
Jul 2026

SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them

SpatialCLI is proposed, a framework that teaches VLMs to reason with spatial tools and progressively internalize the specialist perceptual capabilities they provide and introduces SpatialCLI-Bench, a 516-example benchmark for compositional perception across localization, segmentation, depth, and pose.

Yang Zhou, Zixuan Huang, Sunzhu Li et al. · 0 citations
#small language model Preprint Aug 2026

Retrieval Heads Meet Vision: Uncovering How VLMs Locate and Extract Visual Information

This work introduces Visual Retrieval Heads (VRHs), a small subset of attention heads that are causally responsible for grounding text descriptions to image regions, and shows that scoring attention from output prediction tokens with a sum over the ground-truth referent region most reliably identifies causal heads.

Chanho Park, Daehyeon Choi, Jihyun Lee et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.