Skip to content
Conference

Semantic relevance guided grounding for MLLM-based embodied navigation

Aug 2026 · International Conference on Advanced Sensing and Intelligent Systems · Vol 14309, pp. 1430916 - 1430916-6 · 0 citations · 14 references
Engineering

TL;DR

Experiments on the AI2Thor platform demonstrate that SRG-Nav outperforms baseline methods in both success rate and path efficiency, validating that structured semantic-visual prompts significantly improve the robustness of embodied navigation.

Abstract

Multimodal Large Language Models (MLLMs) based Embodied navigation faces a severe challenge where key cues are easily overwhelmed by complex environmental noise, leading to inefficient decision-making. To address this, we propose a Semantic Relevance Guided grounding enhanced navigation framework(SRG-Nav). The core idea of our approach lies in utilizing semantic relevance to guide visual and language attention. By evaluating the correlation between scene entities and the navigation goal, SRG-Nav adds ranked high-value cues to system prompts and maps them back into the visual space to generate explicit bounding boxes. This mechanism explicitly directs the MLLM to focus on task-relevant entities and regions while effectively suppressing environmental noise. Experiments on the AI2Thor platform demonstrate that SRG-Nav outperforms baseline methods in both success rate and path efficiency, validating that structured semantic-visual prompts significantly improve the robustness of embodied navigation.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

CGFM-Nav: Cognitive Graph-Field Memory for Semantic-Guided Lifelong Multimodal Embodied Navigation

CGFM-Nav is introduced, a foundation-model-based framework for lifelong multimodal navigation that integrates task-relevant subgraph selection, VLM reasoning, and verification feedback into a closed decision loop, and preliminary experiments show that CGFM-Nav improves the overall success rate, demonstrating the effect...

Yu-Xiang Xiao, Xibei Chen, Xin Zhou et al. · 0 citations
#artificial intelligence Preprint Aug 2026

LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation

LightNav-0 is presented, a compact generalist embodied navigation model that elicits the spatial intelligence of a pretrained VLM and aligns it with navigation, without task-specific prediction heads, and establishes compact VLMs as a unified and transferable backbone for generalist embodied navigation.

Shao-An Wang, Ao-Cheng Luo, Fei Huang et al. · 2 citations
Jul 2026

Towards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation

Vision-and-Language Navigation in continuous environments (VLN-CE) requires an agent to ground language in egocentric observations and plan in unseen scenes. Although recent multimodal large models and world-model-based methods have improved navigation, they often preserve excessive task-irrelevant detail, weakening ge...

Yihao Wu, Chen-Yi Xu, Li-Qi Yan et al. · 0 citations
Jul 2026

Explainable and Resource-Efficient Spatial Reasoning in Multimodal LLMs for Decision-Critical Applications

ByDeWay-V2 is proposed, which integrates explicit spatial relational context alongside depth cues, expressed as human-readable predicates that serve as auditable evidence for downstream decision support, showing the framework's suitability for resource-constrained, real-time decision-support settings.

Piyush Jain, Kousik Dasgupta, Rajarshi Roy et al. · 0 citations
Preprint Aug 2026

Embodied Multimodal Grounding for Open-Vocabulary Mobile Manipulation via Semantic 3D Gaussian Splatting

Embodied mobile manipulation requires language, visual observations, three-dimensional scene structure, and action feasibility to be aligned before execution. We study open-vocabulary target grounding with few-shot manipulation in local household workspaces and present an embodied multimodal grounding framework that in...

Huosen Ou, Dong-Ni Song, Yuncong Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.