FSD-VLN is proposed, a fast-slow dual-system architecture disentangling semantic reasoning and low-latency flight command generation that validates the benefit of decoupled semantic-control modeling and provides a practical paradigm for long-horizon aerial VLN.
Abstract
Vision-Language Navigation (VLN) enables UAV autonomous navigation in unknown environments by mapping language instructions to real-time visual inputs. Compared with GPS-dependent or pre-programmed navigation, VLN supports intuitive human-machine interaction and stronger environmental adaptability, requiring tight integration of high-level semantic reasoning and low-latency flight control.Existing methods suffer from structural misalignment between global multimodal understanding and sequential action generation, causing jittery trajectories and severe decision latency for long-horizon aerial navigation. To solve this issue, we propose FSD-VLN, a fast-slow dual-system architecture disentangling semantic reasoning and low-latency flight command generation.The framework has two asynchronous branches: a slow stream extracting stable semantic priors from pre-trained vision-language models, and a Diffusion Transformer (DiT) fast stream modeling cross-temporal action distributions to produce consistent flight outputs. We further introduce a time-aware adaptive optimizer to stabilize long-sequence training and reduce gradient oscillation.Large-scale low-altitude simulation experiments show FSD-VLN achieves up to 2X higher navigation success rates on unseen scenes than SOTA methods, while cutting single-action inference delay and total task runtime by over 50%. Our work validates the benefit of decoupled semantic-control modeling and provides a practical paradigm for long-horizon aerial VLN.
VLN-AVP is proposed, a zero-shot navigation framework for AVP tasks that eliminates the dependency on pre-built maps, interprets semantic environmental contexts in parking scenarios, and enables intuitive navigation following natural language instructions and introduces a hybrid memory system.
Yijiang Li, Xiangru Mu, Changze Li et al.· arXiv.org· 0 citations
Vision-and-Language Navigation (VLN) requires an agent to navigate through unseen 3D environments according to natural-language instructions. Explicit reasoning can improve instruction understanding and semantic grounding, but autoregressive generation at every step accumulates large decision-stage latency over multi-s...
Zhi-Xin Wang, Chengzheyi Yao, Le-Yuan Liu et al.· 0 citations
This work proposes ARIES-Mission2, a zero-shot Vision-Language-Action (VLA) framework that decouples visual-semantic perception from physical route optimization in low-altitude Unmanned Aerial Vehicle (UAV) mission generation and indicates that the TSP module maintains lower growth in computation time as the number of...
Jun-Hao Wei, Yan-Xiao Li, Hao-Chen Li et al.· 1 citation
This work presents an instruction-grounded semantic enhancement module that injects object-level semantics and relative spatial cues into the current observation state, and develops a relevance-aware dynamic temporal aggregation strategy that reweights the full history buffer while converting a few high-relevance frame...
GN (Pangu Navigator), an offline VLN action-prediction system built on OpenPangu-7B, combines mixed-precision computation, selective FP32 computation, and DeepSpeed ZeRO-2 on eight Ascend 910B NPUs and reports metrics quantify offline expert-action alignment rather than closed-loop navigation success.
Li Xian, Mingxi Li, Yizheng Wang et al.· arXiv.org· 0 citations
A simple and effective approach to apply test-time scaling to VLN for UAV navigation through an iterative refinement process that requires no extra model training, guiding the model to re-evaluate its initial navigation plan for better accuracy and safety.