Skip to content

Seeing Is Not Measuring: Tool-Augmented Metric Spatial Reasoning for Vision-Language Models

Sep 2026 · 0 citations · 28 references
Computer Science

TL;DR

A modular, predictor agnostic, tool-augmented framework that equips a small VLM with geometric tools: 3D object detection, metric depth estimation, and deterministic solvers for distance, size and bearing, matching a scripted pipeline on three of four tasks.

Abstract

Vision-Language Models (VLMs) describe scenes well but reason poorly about metric 3D structure such as absolute distances, physical sizes, or egocentric directions. We present a modular, predictor agnostic, tool-augmented framework that equips a small VLM (Qwen3.5-4B) with geometric tools: 3D object detection, metric depth estimation, and deterministic solvers for distance, size and bearing. Each object is detected in the camera frame of its own best view, and the tools use that frame's pose to lift every detection into one shared world frame. Moving metric computation out of the model's weights and into explicit solvers yields large gains on three of four ReVSI-Bench tasks: with a strong monocular detector (WildDet3D), absolute distance rises from 0.46 to 0.74 Mean Relative Accuracy (MRA), relative distance from 39.1% to 67.4%, and relative direction from a below-chance 25.9% to 73.4%. Because any detector can be swapped in behind the tool interface, comparing real detectors against ground-truth boxes separates perception error from reasoning error: orchestration costs only 0.03 MRA. Object size is bounded by the detector: the tools are near-exact on groundtruth boxes (0.97) yet the best real detector barely beats the no-tool baseline (0.61 vs. 0.58), because size reads straight off a box extent monocular detectors get wrong. Without a predefined recipe, the model already sequences the tools correctly on its own, matching a scripted pipeline on three of four tasks.

View source

Similar papers

Preprint Sep 2026

Geometric Encoding for Spatial Reasoning in Vision-Language Models

Vision-Language Models (VLMs) are far more reliable at recognizing what appears in a video than at reasoning about its spatial and temporal properties, such as metric distances, object dimensions, and consistent object identities across frames. We present Geometric Code, a perception-to-geometry pipeline that computes...

Antonio Jun, Hao-Shui Yu, Zheng Lu et al. · 0 citations
Preprint Aug 2026

GaussVLA: Geometry-Aware Spatial Reasoning for Vision-Language-Action Model

GaussVLA is proposed, a Mamba-based VLA that incorporates two custom modules: Gaussian Spatial Tokenizer (GST) to lift frozen semantic and depth features into compact 3D Gaussian tokens, and Depth-Aware Chain-of-Thought (DA-CoT) that performs structured, non-autoregressive geometric reasoning under language and flow-ti...

MD SELIM SAROWAR, Md Tanvir Islam, Sungho Kim et al. · 1 citation
#artificial intelligence Preprint Sep 2026

SceneBench: A Hierarchical Benchmark for Vision-Language Understanding of 3D Scenes

SceneBench is introduced, a benchmark of 966 photorealistic 3D scenes reconstructed with Gaussian Splatting and densely annotated with hierarchical semantics spanning scenes, rooms, functional areas, object groups, and individual objects that provides a realistic testbed for developing and evaluating models capable of...

Anubhav Khanal, Prabigya Acharya, Roshni Poudel et al. · 0 citations
Preprint Aug 2026

RefineAny3D: Depth Refinement as Semantic Alignment for Monocular 3D Detection

This work presents RefineAny3D, a vision-language model that corrects depth without ever predicting a numerical value, and delivers consistent gains across closed-set detectors, open-vocabulary detectors, and 3D auto-labeling tools, and generalizes to novel categories, scenes, and cameras without retraining.

Zhihao Zhang, Geng-Wei Zhang, Tianlong Chen et al. · 0 citations
Preprint Sep 2026

AnchorVLN: Geometry-Anchored Vision-Language Grounding Reasoning for Open-Vocabulary Navigation

Vision-Language Navigation (VLN) in unseen indoor environments is useful in real-world robotics, where an agent must follow natural-language instructions, locate objects, and answer spatial questions without a pre-built map or fixed object vocabulary. Multimodal vision-language models (VLMs) provide strong open-vocabul...

Long Giang Vu, Cheng-Kai Yao, Yu-Xin Liu et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.