Skip to content
Preprint

VLMs Can Describe, But Not Measure: Object-Centric Scene Understanding for Robotic Manipulation

Sep 2026 · 0 citations · 29 references
Computer Science

TL;DR

A VLM-driven, modular perception framework for scene understanding using off-the-shelf approaches that preserves strong semantic performance while substantially improving localization and depth estimation over direct VLM inference.

Abstract

Robotic operation in previously unseen environments requires both semantic understanding and reliable metric information. While vision--language models (VLMs) provide strong semantic capabilities, their geometric estimates remain less reliable. In this paper, we propose a VLM-driven, modular perception framework for scene understanding using off-the-shelf approaches. Starting from a single RGB-D observation, the scene is segmented into object-level regions, annotated by a VLM, and grounded with depth information to construct a task-independent object-centric representation. Experiments on 151 tabletop scenes show that the proposed decomposition preserves strong semantic performance while substantially improving localization and depth estimation over direct VLM inference. The resulting representation is also integrated with a task-planning framework for robotic execution.

View source

Similar papers

May 2025

Object Concepts Emerge from Motion

Object concepts play a foundational role in human visual cognition, enabling perception, memory, and interaction in the physical world. Inspired by findings in developmental neuroscience - where infants are shown to acquire object understanding through observation of motion - we propose a biologically inspired framewor...

Hao Liang, Xiao-Hui Wang, Zhi-Chao Li et al. · 0 citations
Open access Sep 2026

Open-Vocabulary Instance Segmentation for Scene Understanding in Mobile Robots

An open-vocabulary semantic mapping pipeline is presented that integrates TALOS (TAgging–LOcation–Segmentation–Segmentation) with the probabilistic, instance-aware Voxeland framework and shows a better balance between map completeness, geometric clarity, semantic coherence, and instance separation.

Macoris Decena-Gimenez, Pepe Ojeda, J. Ruiz-Sarmiento et al. · 0 citations
Preprint Sep 2026

SAVLA: Symmetry-Aware Vision-Language-Action Models for Robotic Manipulation

Vision-language-action (VLA) models have become the dominant paradigm for language-conditioned robot manipulation. However, although images and language instructions inherently encode geometric information, VLAs acquire their spatial competence purely from demonstrations. As a result, they are reliable only within the...

Jun-Le Li, Weixian Waylon Li, Fu-Xiang Wu et al. · 1 citation
#small language model Preprint Sep 2026

Seeing Is Not Measuring: Tool-Augmented Metric Spatial Reasoning for Vision-Language Models

A modular, predictor agnostic, tool-augmented framework that equips a small VLM with geometric tools: 3D object detection, metric depth estimation, and deterministic solvers for distance, size and bearing, matching a scripted pipeline on three of four tasks.

Kai Glantz, Clemens Grange · 0 citations
Preprint Aug 2026

OptiSight: Bridging Semantic Reasoning and Geometric Control for Embodied Navigation

OptiSight is proposed, a hybrid framework that combines Vision-Language Model reasoning with deterministic visual servoing through a finite-state Chain-of-Thought architecture that demonstrates reliable zero-shot navigation across diverse indoor scenarios, including obstacle avoidance and semantic ambiguity, while oper...

А. А. Аван, Jordi Sanchez-Riera · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.