Skip to content
Preprint

DroneWAM: Efficient World Action Model for Drone Visual Navigation

Sep 2026 · 0 citations · 44 references
Computer Science

TL;DR

DroneWAM adopts a JEPA-based architecture to model future states directly in representation space, avoiding the cost of explicit future image generation and reduces the average prediction depth from 8 to 4.58 while improving trajectory accuracy, demonstrating that predictive computation can be allocated more effectively across scenes.

Abstract

World-action models give visual navigation agents a way to anticipate how candidate actions will change future observations and to act from the predicted consequences. For drones, this capability must operate under tight accuracy and efficiency constraints. We present DroneWAM, an efficient world-action model for drone visual navigation. DroneWAM adopts a JEPA-based architecture to model future states directly in representation space, avoiding the cost of explicit future image generation. A pretrained Resampler further compresses dense encoder features into fewer latent tokens, reducing the computation repeated at each imagined step. We also introduce adaptive rollout, where a preference-trained Gate adaptively allocates prediction depth according to the current scene. To support learning under richer aerial motion, we construct DroneNav-6D, a simulated visual navigation dataset with synchronized RGB observations, 6-DoF flight trajectories, control commands, and randomized wind disturbances. On DroneNav-6D, DroneWAM achieves the best trajectory accuracy among the compared methods. Adaptive rollout further reduces the average prediction depth from 8 to 4.58 while improving trajectory accuracy, demonstrating that predictive computation can be allocated more effectively across scenes. \href{https://github.com/1e12Leon/DroneWAM}{Codes and data} will be released.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Skytopia: Monocular Drone Navigation with Action-Conditioned Latent World Models

Monocular drone navigation requires reaching a goal in an unseen environment from a single forward-facing camera, which offers few cues for depth and scale. World models address this by modelling how observations evolve under actions, but they are built to be executed: the prediction is produced at deployment and fed b...

Yu-Hang Zhang, Rangya Zhang, Yu-Jing Shang et al. · 0 citations
Open access Aug 2026

DevGRU: Depth-Guided Visual Navigation Using a Collision-Aware Recurrent Model

The proposed DevGRU navigation system employs an action predictor that generates collision-aware future trajectories, enabling effective avoidance of immediate obstacles and has a relatively small number of trainable parameters, resulting in the fastest inference time among the baselines.

Kyung Min Han, Eunsom Kim, Young J. Kim · 0 citations
Preprint Sep 2026

V2X-WAM: A Cooperative World Action Model for End-to-End Autonomous Driving

Vehicle-infrastructure cooperation can complement onboard sensing with broader and more informative observations of the traffic environment, providing valuable support for end-to-end autonomous driving. However, existing cooperative driving methods mainly exploit roadside information to enhance the representation of th...

Jun-Wei You, Wei-Zhe Tang, Can Wang et al. · 0 citations
Preprint Sep 2026

VLM-MPPI: Grounding Natural Language in Behaviorally Diverse Trajectories for Aerial Navigation

A hierarchical UAV navigation framework that aligns natural-language intent with dynamically feasible flight behaviors in cluttered indoor environments is presented, and a parallelized ensemble of six behavior-conditioned Model Predictive Path Integral (MPPI) planners are employed.

Hanbing Zhang, Fang-Guo Zhao, Ze-Rui Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

DiffWAM: A Fast and Efficient Navigation World Action Model

Pretrained video foundation models encode rich semantic and spatiotemporal priors for embodied navigation, yet converting these priors into UAV motion typically requires expensive future-video synthesis and geometric reconstruction. We investigate whether the motion implicit in future visual prediction can instead be r...

Morui Zhu, Yu-Ze Wu, Xi-Jie Huang et al. · 0 citations
Preprint Sep 2026

PhysWAM: Physically Consistent World Action Model for Autonomous Driving

PhysWAM is presented, a unified world-action model for autonomous driving that co-denoises multiview video, metric depth, and ego motion within a single flow-matching transformer and Coupled Point Projection is introduced, demonstrating that the geometric relationship between scene depth and ego motion provides a direc...

Dhruv Parikh, Feng-Cheng Yu, Quan-Kai Gao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.