Skip to content

Vision-Language-Motion Maps: An Open-Vocabulary, Uncertainty-Aware, Queryable Motion Attribute for 3D Scene Maps

Jul 2026 · arXiv.org · Vol abs/2607.16173 · 2 citations · ⚡ 1 influential · 21 references
Computer Science

TL;DR

Vision-Language-Motion Maps is introduced, an open-vocabulary, language-queryable 3D map queried through a rule-based intent router over open-vocabulary object nouns, in which each element carries a fused motion attribute: a VLM/LLM semantic movability prior combined with geometrically observed cross-frame motion, together with a per-element uncertainty.

Abstract

Open-vocabulary 3D maps let robots answer language queries about what and where, but they assume a static world and cannot answer queries about how scene elements behave. We introduce Vision-Language-Motion Maps (VLMM), an open-vocabulary, language-queryable 3D map - queried through a rule-based intent router over open-vocabulary object nouns, not a general natural-language interface - in which each element carries a fused motion attribute: a VLM/LLM semantic movability prior combined with geometrically observed cross-frame motion, together with a per-element uncertainty. Queries reduce to attribute filters that distinguish what has been seen to move, what could move but has not, and what stays still. On a controlled simulator benchmark with exact ground truth (AI2-THOR, three scene types) we show through ablation that the schema fields are non-substitutable: a semantic-only baseline fails motion queries even with strong features, and neither motion field substitutes for the other (the prior cannot answer"what is moving,"observed motion cannot answer"what could move"). On real dynamic RGB-D (TUM and Bonn, six sequences) we show the uncertainty channel - our key difference from prior fused-motion work - consistently improves moving-vs-static average precision and reduces false motion flags, and that it is robust to estimated (noisy) poses. The raw confidence is not calibrated, but post-hoc isotonic calibration reaches an expected calibration error of 0.10. VLMM is a representation contribution: the closest prior maps each lack at least one of the four properties - open-vocabulary, language-queryable, fused prior-and-observed motion, and per-element uncertainty - that our combination provides.

View source

Similar papers

Open access Jul 2026

SuperMap: A Spatio-Temporal SLAM System for Visual-Language Navigation

This work presents SuperMap, a 4D spatio-temporal mapping framework for language-guided navigation that integrates high-frequency geometric SLAM with asynchronous open-vocabulary perception and releases the full system as open-source to provide the community with a deployable baseline for open-vocabulary spatio-tempora...

Shibo Zhao, Guofei Chen, Honghao Zhu et al. · 2 citations
Jul 2026

Memory for Attention: Language-Conditioned Re-Perception with a Vision-Language-Motion Map

A persistent map's memory yields the best schedule (held-out), matching an oracle, while the memoryless VLM prior is poor, showing a persistent map's memory earns its keep not by telling the robot how to walk around a room, but by telling it what to pay attention to.

Dibyendu Ghosh · 0 citations
Preprint Aug 2026

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models

This work introduces SPATIALQUERY, a training- free framework for CIDQ reasoning from a single RGB image, together with SPATIALQUERY-1M, a benchmark containing over one million RGB-only question-answer pairs from 200 indoor scenes, and proposes Uncertainty-Aware Chain-of-Thought (UA-CoT) prompting, which incorporates g...

Hai-Tra Nguyen, Tung Vu, Cong Tran · 0 citations
Preprint Aug 2026

WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN

WNM-3D, a generative World Navigation Model with 3D scene conditioning for continuous VLN, is presented, showing that WNM-3D outperforms strong VLM-based navigation policies and its 2D-conditioned counterpart in closed-loop navigation.

Yue-Hao Huang, Yunzi Wu, Xiaotao Zhang et al. · 1 citation
Preprint Aug 2026

SAP-Nav: Spatial Semantic Representation Meets Active Perception for Hierarchical Open-Vocabulary Object Navigation

SAP-Nav is presented, a fully online, zero-shot framework that supports both hierarchical and standard category-level OVON without task-specific training or precomputed scene maps and is designed for hierarchical OVON without task-specific training or precomputed scene maps.

Xuetong Pei, Jian Liu, Vidura Munasinghe et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.