This work proposes an object--path graph that unifies open-vocabulary semantic reasoning with topological navigation, and introduces a navigation strategy that combines global path planning with local inter-node execution through lightweight node localization and semantic visual servoing, enabling navigation directly over the graph without dense metric reconstruction.
Abstract
Vision-language navigation requires embodied agents to navigate environments using natural language instructions and visual observations. Existing approaches typically decompose navigation into sequential language-guided decisions or rely on online exploration without prior environmental knowledge. Scene graph representations offer compact semantic memory but remain decoupled from downstream navigation, which still depends on dense metric maps. To close this gap, we propose an object--path graph that unifies open-vocabulary semantic reasoning with topological navigation. The proposed representation jointly supports semantic grounding, graph-based localization, and navigation within a single lightweight topological framework. Building on this graph, we introduce a navigation strategy that combines global path planning with local inter-node execution through lightweight node localization and semantic visual servoing, enabling navigation directly over the graph without dense metric reconstruction. Experiments on HM3D and Replica demonstrate competitive performance in open-vocabulary object grounding through the proposed hierarchical graph structure, while achieving effective navigation performance. Real-world robot experiments further validate the practicality of the proposed framework.
CGFM-Nav is introduced, a foundation-model-based framework for lifelong multimodal navigation that integrates task-relevant subgraph selection, VLM reasoning, and verification feedback into a closed decision loop, and preliminary experiments show that CGFM-Nav improves the overall success rate, demonstrating the effect...
Yu-Xiang Xiao, Xibei Chen, Xin Zhou et al.· 0 citations
Large language models (LLMs) are increasingly used as natural-language interfaces for robotic systems, yet their integration with Robot Operating System (ROS)-based navigation remains limited by two gaps. First, navigation data such as occupancy grids are represented as raw geometric messages that are difficult for LLM...
Jungsoo Lee, Jaegyun Park, Wansoo Kim· 0 citations
Map-based vision-language navigation (VLN) relies on persistent spatial representations to connect language understanding with geometric planning. However, acquiring semantics beyond the needs of the current instruction can introduce unnecessary perception cost and irrelevant annotations. Continuously accumulating unre...
Quan-Hua Chen, Ju-Han Kang, Run-Feng Lin et al.· 0 citations
Vision-Language Navigation (VLN) in unseen indoor environments is useful in real-world robotics, where an agent must follow natural-language instructions, locate objects, and answer spatial questions without a pre-built map or fixed object vocabulary. Multimodal vision-language models (VLMs) provide strong open-vocabul...
Long Giang Vu, Cheng-Kai Yao, Yu-Xin Liu et al.· 0 citations
Continuous-environment vision-and-language navigation (VLN-CE) requires interpreting natural-language instructions in unseen 3D environments and executing continuous low-level actions. Existing methods often depend on LiDAR, panoramic cameras, or extra sensors; separate geometric-mapping and semantic-navigation visual...
Jian-He Zhao, Yan-Hua Qiu, Zhi-Yu Zhang et al.· 0 citations
SAP-Nav is presented, a fully online, zero-shot framework that supports both hierarchical and standard category-level OVON without task-specific training or precomputed scene maps and is designed for hierarchical OVON without task-specific training or precomputed scene maps.
Xuetong Pei, Jian Liu, Vidura Munasinghe et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.