Skip to content

DA-Nav: Direction-Aware City-Scale Vision-Language Navigation

Jul 2026 · arXiv.org · Vol abs/2607.11638 · 0 citations · 37 references
Computer Science

TL;DR

DA-Nav is proposed, a Direction-Aware vision-language Navigation framework that reformulates navigation as a discrete spatial grounding problem on the egocentric 2D image plane, outperforming existing State-of-The-Art (SoTA) methods while maintaining a substantially stronger recovery capability.

Abstract

City-scale outdoor navigation is currently hindered by the heavy reliance on dense maps or costly navigation supervision. In this work, we introduce a novel paradigm for leveraging directional instructions from commercial navigation tools (e.g., Google Maps). To bridge the gap between commercial instructions and executable navigation actions, while mitigating long-horizon error accumulation through robust trajectory recovery, we propose DA-Nav, a Direction-Aware vision-language Navigation framework that reformulates navigation as a discrete spatial grounding problem on the egocentric 2D image plane. To achieve trajectory recovery, DA-Nav employs a Chain-of-Thought (CoT) reasoning process encompassing deviation assessment, action prediction, and target grid selection. We further introduce ReDA, a dataset that provides direction-aware instructions and recovery trajectories to enhance spatial grounding and support CoT recovery reasoning. Extensive experiments in CARLA demonstrate that DA-Nav achieves a high success rate of 56.16% in unseen urban environments, outperforming existing State-of-The-Art (SoTA) methods while maintaining a substantially stronger recovery capability. Furthermore, without fine-tuning, DA-Nav seamlessly adapts to both quadruped and humanoid robots, enabling stable kilometer-scale closed-loop outdoor navigation in complex real world environments.

View source

Similar papers

Preprint Aug 2026

From Semantic Grounding to Decision Optimization: A Unified Framework for Long-Horizon UAV Vision-Language Navigation

This work presents an instruction-grounded semantic enhancement module that injects object-level semantics and relative spatial cues into the current observation state, and develops a relevance-aware dynamic temporal aggregation strategy that reweights the full history buffer while converting a few high-relevance frame...

Ze-Yuan Ma, Jiaxin Chen, Di Huang · 0 citations
Jul 2026

VoLN: Vision-Only Long-Horizon Navigation - Paradigm, Benchmark, and Method

This work instantiates VoLN for aerial navigation through VoLN-UAV, a 7,210-episode benchmark that combines long-horizon goal-directed flight, continuous 3D motion, large viewpoint changes, and context-dependent beacon selection and reveals substantial remaining challenges in long-horizon evidence integration, cross-vi...

Jiabin Lou, Hao-Peng Wang, Yuan-Shuai Wang et al. · 0 citations
Preprint Sep 2026

Map the Possibilities: Spatial Belief Fields for Language-Goal Aerial Navigation

Language-goal aerial navigation requires an agent to local- ize a potentially unobserved target from relational instruc- tions and partial observations, and translate this inference into metric actions in large-scale continuous environments. Existing methods often reduce language grounding to one single waypoint or act...

Hao-Tian Xu, Yue Hu, Zheng-Qiu Zhu et al. · 0 citations
Preprint Aug 2026

OccPlanner: Goal-Aware Occupancy-Conditioned Diffusion Planner for PixelGoal Navigation

PixelGoal navigation specifies targets directly in the agent's camera view, providing a natural interface between high-level visual reasoning and low-level navigation. Depth can lift a visible target pixel into a metric PointGoal, but this estimate becomes unreliable under occlusion or sensor noise. Moreover, a PointGo...

Binling Huang, Nian-Jin Ye, Xi Yang et al. · 0 citations
Open access Aug 2026

DevGRU: Depth-Guided Visual Navigation Using a Collision-Aware Recurrent Model

The proposed DevGRU navigation system employs an action predictor that generates collision-aware future trajectories, enabling effective avoidance of immediate obstacles and has a relatively small number of trainable parameters, resulting in the fastest inference time among the baselines.

Kyung Min Han, Eunsom Kim, Young J. Kim · 0 citations
Conference Open access Sep 2026

FILD-Nav:Vision-and-Language Navigation with Instruction Landmark Features in Continuous Environments

FILD-Nav extracts task-relevant landmarks from instructions and incorporates landmark semantics into both waypoint prediction and topological planning, which improves waypoint relevance and enables more effective long-horizon navigation.

Chuan-Ye Hu, Lu-Lu Liu, Huai-Wei Si et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.