Jul 2026
When Depth Is Better Told Than Shown: Depth-Ordinal Prompting for Vision-Language Spatial Reasoning
This work proposes Depth-Ordinal Prompting (DOP), a training-free method that converts monocular depth into a single question-targeted ordinal text cue at the queried objects, without adding a depth image, training a module, injecting features, or using labels.
Quynh T. N. Vo, Phuc T. Dao, Cong-Duy Nguyen et al.
· arXiv.org · 0 citations