A VLM-driven, modular perception framework for scene understanding using off-the-shelf approaches that preserves strong semantic performance while substantially improving localization and depth estimation over direct VLM inference.
Abstract
Robotic operation in previously unseen environments requires both semantic understanding and reliable metric information. While vision--language models (VLMs) provide strong semantic capabilities, their geometric estimates remain less reliable. In this paper, we propose a VLM-driven, modular perception framework for scene understanding using off-the-shelf approaches. Starting from a single RGB-D observation, the scene is segmented into object-level regions, annotated by a VLM, and grounded with depth information to construct a task-independent object-centric representation. Experiments on 151 tabletop scenes show that the proposed decomposition preserves strong semantic performance while substantially improving localization and depth estimation over direct VLM inference. The resulting representation is also integrated with a task-planning framework for robotic execution.
RECAST is a robot navigation framework that combines the reasoning of a VLM with the spatial grounding of vision foundation models to build an Actionable Cost map and improves success over the strongest prior method.
Incheol Cho, Jintae Park, Jinkyu Kim et al.· 0 citations
Object concepts play a foundational role in human visual cognition, enabling perception, memory, and interaction in the physical world. Inspired by findings in developmental neuroscience - where infants are shown to acquire object understanding through observation of motion - we propose a biologically inspired framewor...
Hao Liang, Xiao-Hui Wang, Zhi-Chao Li et al.· Neural Information Processin...· 0 citations
An open-vocabulary semantic mapping pipeline is presented that integrates TALOS (TAgging–LOcation–Segmentation–Segmentation) with the probabilistic, instance-aware Voxeland framework and shows a better balance between map completeness, geometric clarity, semantic coherence, and instance separation.
Macoris Decena-Gimenez, Pepe Ojeda, J. Ruiz-Sarmiento et al.· Robotics· 0 citations
Vision-language-action (VLA) models have become the dominant paradigm for language-conditioned robot manipulation. However, although images and language instructions inherently encode geometric information, VLAs acquire their spatial competence purely from demonstrations. As a result, they are reliable only within the...
A modular, predictor agnostic, tool-augmented framework that equips a small VLM with geometric tools: 3D object detection, metric depth estimation, and deterministic solvers for distance, size and bearing, matching a scripted pipeline on three of four tasks.
OptiSight is proposed, a hybrid framework that combines Vision-Language Model reasoning with deterministic visual servoing through a finite-state Chain-of-Thought architecture that demonstrates reliable zero-shot navigation across diverse indoor scenarios, including obstacle avoidance and semantic ambiguity, while oper...
А. А. Аван, Jordi Sanchez-Riera· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.