Vision-Language-Motion Maps is introduced, an open-vocabulary, language-queryable 3D map queried through a rule-based intent router over open-vocabulary object nouns, in which each element carries a fused motion attribute: a VLM/LLM semantic movability prior combined with geometrically observed cross-frame motion, together with a per-element uncertainty.
Abstract
Open-vocabulary 3D maps let robots answer language queries about what and where, but they assume a static world and cannot answer queries about how scene elements behave. We introduce Vision-Language-Motion Maps (VLMM), an open-vocabulary, language-queryable 3D map - queried through a rule-based intent router over open-vocabulary object nouns, not a general natural-language interface - in which each element carries a fused motion attribute: a VLM/LLM semantic movability prior combined with geometrically observed cross-frame motion, together with a per-element uncertainty. Queries reduce to attribute filters that distinguish what has been seen to move, what could move but has not, and what stays still. On a controlled simulator benchmark with exact ground truth (AI2-THOR, three scene types) we show through ablation that the schema fields are non-substitutable: a semantic-only baseline fails motion queries even with strong features, and neither motion field substitutes for the other (the prior cannot answer"what is moving,"observed motion cannot answer"what could move"). On real dynamic RGB-D (TUM and Bonn, six sequences) we show the uncertainty channel - our key difference from prior fused-motion work - consistently improves moving-vs-static average precision and reduces false motion flags, and that it is robust to estimated (noisy) poses. The raw confidence is not calibrated, but post-hoc isotonic calibration reaches an expected calibration error of 0.10. VLMM is a representation contribution: the closest prior maps each lack at least one of the four properties - open-vocabulary, language-queryable, fused prior-and-observed motion, and per-element uncertainty - that our combination provides.
This work presents SuperMap, a 4D spatio-temporal mapping framework for language-guided navigation that integrates high-frequency geometric SLAM with asynchronous open-vocabulary perception and releases the full system as open-source to provide the community with a deployable baseline for open-vocabulary spatio-tempora...
A persistent map's memory yields the best schedule (held-out), matching an oracle, while the memoryless VLM prior is poor, showing a persistent map's memory earns its keep not by telling the robot how to walk around a room, but by telling it what to pay attention to.
This work introduces SPATIALQUERY, a training- free framework for CIDQ reasoning from a single RGB image, together with SPATIALQUERY-1M, a benchmark containing over one million RGB-only question-answer pairs from 200 indoor scenes, and proposes Uncertainty-Aware Chain-of-Thought (UA-CoT) prompting, which incorporates g...
WNM-3D, a generative World Navigation Model with 3D scene conditioning for continuous VLN, is presented, showing that WNM-3D outperforms strong VLM-based navigation policies and its 2D-conditioned counterpart in closed-loop navigation.
Yue-Hao Huang, Yunzi Wu, Xiaotao Zhang et al.· 1 citation
Results indicate that explicit, refreshable 3D semantic grounding can improve robustness under clutter, occlusion, viewpoint variation, and embodiment constraints.
Huosen Ou, Dong-Ni Song, Yuncong Wang et al.· 0 citations
SAP-Nav is presented, a fully online, zero-shot framework that supports both hierarchical and standard category-level OVON without task-specific training or precomputed scene maps and is designed for hierarchical OVON without task-specific training or precomputed scene maps.
Xuetong Pei, Jian Liu, Vidura Munasinghe et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.