Sep 2026· Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence· pp. 7566-7574· 0 citations· 31 references
TL;DR
FILD-Nav extracts task-relevant landmarks from instructions and incorporates landmark semantics into both waypoint prediction and topological planning, which improves waypoint relevance and enables more effective long-horizon navigation.
Abstract
Vision-and-language navigation (VLN) requires agents to follow natural language instructions to navigate autonomously in continuous environments. However, existing approaches often lack high-level semantic guidance in waypoint prediction and explicit language–landmark alignment in cross-modal planning. To address these limitations, we propose FILD-Nav, a vision-and-language navigation framework that integrates instruction landmark features. FILD-Nav extracts task-relevant landmarks from instructions and incorporates landmark semantics into both waypoint prediction and topological planning. Specifically, landmark-guided waypoint prediction improves waypoint relevance, while landmark-enhanced cross-modal planning enables more effective long-horizon navigation. Extensive experiments on the VLN-CE benchmark demonstrate that FILD-Nav consistently outperforms prior methods, achieving improvements of 2% in Success Rate (SR), 3% in Success weighted by Path Length (SPL), and 7% in Oracle Success Rate (OSR), particularly in unseen environments.
Vision-and-Language Navigation (VLN) tasks require an agent to interpret natural language instructions and visual observations to navigate complex environments. Existing methods mostly construct topological or semantic maps and rely on the Large Language Model (LLM) for navigation decision-making, however, they still suffer from limited adaptability to dynamic environments and robustness to complex instructions, with task performance being limited by LLM performance. To overcome these limitations, we introduce the Real-time Occupancy-aware Visual Language Map (RO-VLMap). This framework quantifies the instantaneous risks posed by moving entities in the environment by fusing real-time occupancy sensing with visual language 3D reconstruction. Specifically, RO-VLMap, when combined with our proposed Navigation Adaptive Module (NAM), leverages the complementary advantages of the Knowledge Graph (KG) and LLM to parse open-vocabulary instructions into precise navigation goals. By planning over a unified occupancy-aware map, the agent proactively generates safe paths that avoid dynamic obstacles. Experimental results show that RO-VLMap significantly improves navigation success rates and efficiency. Furthermore, it demonstrates strong robustness in unseen scenarios, providing a practical solution for autonomous navigation of embodied agents in complex real-world environments.
Yuan Liu, Chuang Hu, Nan Ding· International Conferences on...· 0 citations
This work instantiates VoLN for aerial navigation through VoLN-UAV, a 7,210-episode benchmark that combines long-horizon goal-directed flight, continuous 3D motion, large viewpoint changes, and context-dependent beacon selection and reveals substantial remaining challenges in long-horizon evidence integration, cross-view goal matching, and closed-loop stability.
Jiabin Lou, Hao-Peng Wang, Yuan-Shuai Wang et al.· arXiv.org· 0 citations
Vision-and-Language Navigation (VLN) requires an agent to navigate through unseen 3D environments according to natural-language instructions. Explicit reasoning can improve instruction understanding and semantic grounding, but autoregressive generation at every step accumulates large decision-stage latency over multi-step navigation. We propose VerNav, a verifier-first framework for low-latency LLM-based VLN. The verifier reduces decision-stage latency by replacing per-step autoregressive generation with batched action verification, while an entropy-based adaptive generator is invoked only for uncertain decisions to produce compact state evidence. To further improve navigation performance with the verifier, we introduce a two-stage alignment scheme: (i) VPO improves local action-preference alignment in static verifier training, and (ii) step-level reinforcement fine-tuning provides dense progress rewards over multi-step navigation rollouts during dynamic task execution. Experiments on the Room-to-Room (R2R) benchmark show that the verifier-only decision path of VerNav achieves competitive navigation performance among representative LLM-based VLN agents while reducing average decision-stage LLM latency per step by more than $10\times$ compared with autoregressive methods.
Zhi-Xin Wang, Chengzheyi Yao, Le-Yuan Liu et al.· 0 citations
This survey revisits VLN from a component-internal perspective, viewing it as a navigation system composed of internal components such as instructions, environment representations, and embodied agents, and organizing existing work around the functions and interactions of these components.
This paper presents the first systematic benchmark of 17 edge-deployable SLMs against 4 online APIs for robotic navigation instruction decomposition, and proposes a lightweight hybrid semantic-geometric goal localization framework that combines open-vocabulary object detection, prompted segmentation, and LiDAR geometry to estimate metric goals.
Ali Salmasi, Xian-Jia Yu, Tomi Westerlund· arXiv.org· 0 citations
Vision-Language Navigation in Continuous Environments (VLN-CE) requires an embodied agent to execute long-horizon instructions in unknown environments. Existing zero-shot VLN-CE systems typically maintain spatial states through geometric localization or coordinate-based representations. Recent geometry-constrained navigation removes depth and globally consistent coordinates, but maintaining persistent spatial awareness for place confirmation, progress verification, and recovery remains challenging. We present Navi-Agent, a zero-shot VLN-CE agent that constructs a coordinate-free spatial state from visual observations and executed motion histories. Navi-Agent organizes this state as a navigation topology, where nodes represent visual places and edges represent motion transitions. This representation enables observation-based approximate self-localization, task progress verification, and visual revisitation-based recovery. Navi-Agent performs closed-loop navigation by decomposing instructions into sub-goals, executing local visual navigation, and verifying visited places through the constructed spatial state. Experiments on zero-shot VLN-CE benchmark and real-world robot platforms show that Navi-Agent achieves state-of-the-art performance among geometry-constrained methods while remaining competitive with approaches relying on geometric localization.
Wen-Yuan Xie, Meng-Yang Hong, Yong-Zhong Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.