Vision-language models (VLMs) show strong visual understanding for aerial navigation, but their action generation remains unreliable. We study this gap in the context of orbit-search-based navigation, where an aerial agent circles a reference landmark while a VLM detects a language-described goal. We find that the VLM recognizes goals with stable accuracy across configurations, yet the subsequent approach phase—where the VLM must navigate toward the recognized goal—succeeds only 11.3% of the time. The failure is not a detection problem but a navigation problem: the visual approach prompt uses a format absent from the training data, and no amount of prompt engineering restores performance. We propose Geometric Waypoint Navigation, which eliminates VLM-based navigation after detection. The aerial agent moves toward the detected direction using the same deterministic waypoint-following controller already proven in the orbit phase, requiring no additional training or parameters. On the CityNav test unseen split, success rate increases from 35.5% to 49.0%, substantially surpassing all published baselines including GeoNav (25.9%), HTNav (21.7%), and FlightGPT (21.2%). Detection rates remain exactly constant at 31.0%, confirming that gains originate entirely from improved navigation rather than improved perception. Controlled experiments show that geometric waypoint following outperforms VLM Replay by 8.0 percentage points, and a fixed 30m approach distance outperforms VLM-estimated adaptive distances by 3.0 points. Our findings suggest a broader design principle: when a learned model serves as a detector in an embodied pipeline, separating perceptual decisions from motor execution can substantially improve reliability.
Hao-Tian Xu, Chen-Xu Wang, De-Jun Chen et al.· 2026 12th International Conf...· 0 citations
Language-goal aerial navigation requires an agent to local- ize a potentially unobserved target from relational instruc- tions and partial observations, and translate this inference into metric actions in large-scale continuous environments. Existing methods often reduce language grounding to one single waypoint or action, prematurely collapsing the spatial uncertainty inherent in incomplete evidence and ambiguous relations. To address this limitation, we introduce SBFNav, a closed-loop navigation framework centered on a language- conditioned Spatial Belief Field (SBF). Unlike ego-centric maps that primarily record what has been observed, SBF rep- resents a task-conditioned distribution over plausible target locations, preserving multiple spatial hypotheses under par- tial evidence. At each step, this distribution is updated from accumulated observations as new evidence becomes avail- able. Built on this representation, SBFNav selects the goal that best aligns with the instruction and observations as a met- ric waypoint for control. Experiments on both the original and revised CityNav benchmarks achieve the best reported overall performance. On the Test Unseen split, our method improves SR from 25.91% to 32.29% and SPL from 19.63% to 30.43%. Ablation studies further confirm the advantages of spatial-belief modeling over single-point prediction.
Hao-Tian Xu, Yue Hu, Zheng-Qiu Zhu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.