Skip to content
Preprint

ForeFly: A Dual-Horizon World Action Model for Aerial Vision-Language Navigation

Sep 2026 · 0 citations · 52 references
Computer Science

TL;DR

Foresight-Guided Action Refinement (FGAR), which asymmetrically exploits proximal foresight for local action enhancement and route-critical foresight for feature-wise correction and route-level guidance, is introduced, validating the effectiveness of dual-horizon foresight and FGAR learning.

Abstract

Aerial Vision-Language Navigation (AVLN) requires UAVs to maintain reliable instruction following over long trajectories in complex 3D environments. However, existing AVLN approaches are predominantly reactive or limited to single-horizon prediction, overlooking complementary future cues across different temporal horizons. To address this limitation, we propose ForeFly, a dual-horizon latent world action model that predicts both a proximal future for local continuity and an adaptive route-critical future for long-range guidance. Horizon-specific foresight queries are primed with recent and route-critical visual memories, providing history-aware context for future prediction. To exploit their distinct roles in action generation, we introduce Foresight-Guided Action Refinement (FGAR), which asymmetrically exploits proximal foresight for local action enhancement and route-critical foresight for feature-wise correction and route-level guidance. Experiments on the TravelUAV and UAV-ON benchmarks show that ForeFly consistently outperforms strong baselines across seen and unseen settings, validating the effectiveness of dual-horizon foresight and FGAR learning. The code is available at: https://github.com/kunhuiW/ForeFly

View source

Similar papers

Preprint Aug 2026

DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation

Aerial vision-language navigation (VLN) requires an embodied agent to integrate visual evidence over time, plan future actions, and determine when it has reached a navigation goal under partial observability. Although recent VLA models offer a promising perception-to-action paradigm, adapting them to aerial navigation...

Yan Deng, Fei Xu · 1 citation · ⚡1
Preprint Aug 2026

RefineFly: Failure-Aware Post-Training for Aerial Vision-Language Navigation

RefineFly, a failure-aware RL post-training framework for end-to-end UAV-VLA policies, built on token-level proximal policy optimization (PPO), maintains dynamic failure memory within a two-stage scene curriculum to sustain learning from unresolved tasks as training shifts from empirical to balanced scene sampling.

Boxiong Wang, Hui Kang, Geng Sun et al. · 0 citations
Preprint Aug 2026

From Semantic Grounding to Decision Optimization: A Unified Framework for Long-Horizon UAV Vision-Language Navigation

This work presents an instruction-grounded semantic enhancement module that injects object-level semantics and relative spatial cues into the current observation state, and develops a relevance-aware dynamic temporal aggregation strategy that reweights the full history buffer while converting a few high-relevance frame...

Ze-Yuan Ma, Jiaxin Chen, Di Huang · 1 citation
Preprint Sep 2026

VLM-MPPI: Grounding Natural Language in Behaviorally Diverse Trajectories for Aerial Navigation

A hierarchical UAV navigation framework that aligns natural-language intent with dynamically feasible flight behaviors in cluttered indoor environments is presented, and a parallelized ensemble of six behavior-conditioned Model Predictive Path Integral (MPPI) planners are employed.

Hanbing Zhang, Fang-Guo Zhao, Ze-Rui Li et al. · 0 citations
Preprint Aug 2026

AirForesight: Current-to-Future Spatial Map Imagination with Cross-Space Planning Consistency for UAV-VLN

To make spatial imagination more relevant to navigation, this work introduces a cross-space planning consistency loss that encourages directional agreement between the predicted map-space trajectory and the expert action direction derived from the ground-truth waypoint displacement.

Yutong Liu, Xiao-Jie Li, Ming-Zhu Xu et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.