Foresight-Guided Action Refinement (FGAR), which asymmetrically exploits proximal foresight for local action enhancement and route-critical foresight for feature-wise correction and route-level guidance, is introduced, validating the effectiveness of dual-horizon foresight and FGAR learning.
Abstract
Aerial Vision-Language Navigation (AVLN) requires UAVs to maintain reliable instruction following over long trajectories in complex 3D environments. However, existing AVLN approaches are predominantly reactive or limited to single-horizon prediction, overlooking complementary future cues across different temporal horizons. To address this limitation, we propose ForeFly, a dual-horizon latent world action model that predicts both a proximal future for local continuity and an adaptive route-critical future for long-range guidance. Horizon-specific foresight queries are primed with recent and route-critical visual memories, providing history-aware context for future prediction. To exploit their distinct roles in action generation, we introduce Foresight-Guided Action Refinement (FGAR), which asymmetrically exploits proximal foresight for local action enhancement and route-critical foresight for feature-wise correction and route-level guidance. Experiments on the TravelUAV and UAV-ON benchmarks show that ForeFly consistently outperforms strong baselines across seen and unseen settings, validating the effectiveness of dual-horizon foresight and FGAR learning. The code is available at: https://github.com/kunhuiW/ForeFly
Aerial vision-language navigation (VLN) requires an embodied agent to integrate visual evidence over time, plan future actions, and determine when it has reached a navigation goal under partial observability. Although recent VLA models offer a promising perception-to-action paradigm, adapting them to aerial navigation...
RefineFly, a failure-aware RL post-training framework for end-to-end UAV-VLA policies, built on token-level proximal policy optimization (PPO), maintains dynamic failure memory within a two-stage scene curriculum to sustain learning from unresolved tasks as training shifts from empirical to balanced scene sampling.
Boxiong Wang, Hui Kang, Geng Sun et al.· 0 citations
This work presents an instruction-grounded semantic enhancement module that injects object-level semantics and relative spatial cues into the current observation state, and develops a relevance-aware dynamic temporal aggregation strategy that reweights the full history buffer while converting a few high-relevance frame...
A hierarchical UAV navigation framework that aligns natural-language intent with dynamically feasible flight behaviors in cluttered indoor environments is presented, and a parallelized ensemble of six behavior-conditioned Model Predictive Path Integral (MPPI) planners are employed.
Hanbing Zhang, Fang-Guo Zhao, Ze-Rui Li et al.· 0 citations
To make spatial imagination more relevant to navigation, this work introduces a cross-space planning consistency loss that encourages directional agreement between the predicted map-space trajectory and the expert action direction derived from the ground-truth waypoint displacement.
Yutong Liu, Xiao-Jie Li, Ming-Zhu Xu et al.· 1 citation
This work introduces RiverVLN, to its knowledge the first benchmark designed for long-horizon USV VLN under continuous riverine motion, and PGT-NAV, a phase-grounded temporal navigation framework for USVs.
Jie-Ling Wu, Yue-Hao Huang, Jia-Jun Lv et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.