Aug 2026· International Conference on Advanced Sensing and Intelligent Systems· Vol 14309, pp. 1430916 - 1430916-6· 0 citations· 14 references
Engineering
TL;DR
Experiments on the AI2Thor platform demonstrate that SRG-Nav outperforms baseline methods in both success rate and path efficiency, validating that structured semantic-visual prompts significantly improve the robustness of embodied navigation.
Abstract
Multimodal Large Language Models (MLLMs) based Embodied navigation faces a severe challenge where key cues are easily overwhelmed by complex environmental noise, leading to inefficient decision-making. To address this, we propose a Semantic Relevance Guided grounding enhanced navigation framework(SRG-Nav). The core idea of our approach lies in utilizing semantic relevance to guide visual and language attention. By evaluating the correlation between scene entities and the navigation goal, SRG-Nav adds ranked high-value cues to system prompts and maps them back into the visual space to generate explicit bounding boxes. This mechanism explicitly directs the MLLM to focus on task-relevant entities and regions while effectively suppressing environmental noise. Experiments on the AI2Thor platform demonstrate that SRG-Nav outperforms baseline methods in both success rate and path efficiency, validating that structured semantic-visual prompts significantly improve the robustness of embodied navigation.
CGFM-Nav is introduced, a foundation-model-based framework for lifelong multimodal navigation that integrates task-relevant subgraph selection, VLM reasoning, and verification feedback into a closed decision loop, and preliminary experiments show that CGFM-Nav improves the overall success rate, demonstrating the effect...
Yu-Xiang Xiao, Xibei Chen, Xin Zhou et al.· 0 citations
LightNav-0 is presented, a compact generalist embodied navigation model that elicits the spatial intelligence of a pretrained VLM and aligns it with navigation, without task-specific prediction heads, and establishes compact VLMs as a unified and transferable backbone for generalist embodied navigation.
Shao-An Wang, Ao-Cheng Luo, Fei Huang et al.· 2 citations
Experimental results on REVERIE and SOON datasets demonstrate that ViSMoE outperforms the previous state-of-the-art methods, showing the superiority of the proposed method.
Shuo Feng, Pi-Ji Li· International Conference on...· 0 citations
Vision-and-Language Navigation in continuous environments (VLN-CE) requires an agent to ground language in egocentric observations and plan in unseen scenes. Although recent multimodal large models and world-model-based methods have improved navigation, they often preserve excessive task-irrelevant detail, weakening ge...
Yihao Wu, Chen-Yi Xu, Li-Qi Yan et al.· arXiv.org· 0 citations
ByDeWay-V2 is proposed, which integrates explicit spatial relational context alongside depth cues, expressed as human-readable predicates that serve as auditable evidence for downstream decision support, showing the framework's suitability for resource-constrained, real-time decision-support settings.
Piyush Jain, Kousik Dasgupta, Rajarshi Roy et al.· arXiv.org· 0 citations
Embodied mobile manipulation requires language, visual observations, three-dimensional scene structure, and action feasibility to be aligned before execution. We study open-vocabulary target grounding with few-shot manipulation in local household workspaces and present an embodied multimodal grounding framework that in...
Huosen Ou, Dong-Ni Song, Yuncong Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.