Skip to content
Preprint

EgoPathBench: Evaluating Zero-Shot Egocentric Waypoint Decision-Making in Vision-Language Models

Sep 2026 · 0 citations · 27 references
Computer Science

TL;DR

EgoPathBench, a dataset and five-task benchmark for first-person waypoint decision-making, and Fine-tuning Qwen 3.5 4B on the released training split improves all four reported evaluations across three external spatial benchmarks, with gains of 1.4--9.6 points.

Abstract

Zero-shot waypoint navigation requires vision-language models to select, from the current first-person observation, a sequence of spatial actions that is feasible for the agent and reaches the goal, placing joint demands on the integrated spatial intelligence of today's foundation VLMs. Existing spatial-intelligence benchmarks primarily evaluate isolated judgments of relations, directions, or targets and therefore do not directly measure the integrated navigation ability required to combine target recognition, action-consequence assessment, distance estimation, and path planning. To fill this evaluation gap, we introduce EgoPathBench, a dataset and five-task benchmark for first-person waypoint decision-making. Each question presents an egocentric RGB image, a natural-language goal, and numbered visible waypoints; a model returns traversable candidates or an ordered route. Predictions are evaluated for candidate feasibility, adjacent-edge legality, and goal arrival under point-agent or embodied geometry. EgoPathBench contains 31,852 training, 1,345 validation, and 1,111 benchmark questions and retains at least one geometrically verified reference route for every route question. Across nine VLMs, the highest EgoPath Score is only 28.3. The top-ranked model reaches 35.9% success on Point Path, but only 2.9% and 4.0% on Embodied Path and Intent Path, respectively, showing that current models remain limited in forming complete, goal-consistent routes under embodiment constraints. Beyond the evaluation data, we release the corresponding training resource. Fine-tuning Qwen 3.5 4B on the released training split raises its EgoPath Score from 3.9 to 38.9 and improves all four reported evaluations across three external spatial benchmarks, with gains of 1.4--9.6 points.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

One MLLM, One Call: Efficient Zero-Shot Vision-and-Language Navigation via Spatial-Aware Waypoints

Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires an embodied agent to navigate unseen environments by following natural language instructions. Current zero-shot VLN-CE methods either rely on pre-trained waypoint predictors or require multiple queries to large models per step. To address prohi...

Shi-Qi Pan, Qi Zheng, Hanmeng Sun et al. · 1 citation
Preprint Aug 2026

OccPlanner: Goal-Aware Occupancy-Conditioned Diffusion Planner for PixelGoal Navigation

PixelGoal navigation specifies targets directly in the agent's camera view, providing a natural interface between high-level visual reasoning and low-level navigation. Depth can lift a visible target pixel into a metric PointGoal, but this estimate becomes unreliable under occlusion or sensor noise. Moreover, a PointGo...

Binling Huang, Nian-Jin Ye, Xi Yang et al. · 0 citations
Preprint Sep 2026

FINE: Future-Informed Navigation Encoding for Data-Efficient Vision-Language Navigation

Adapting vision-language navigation (VLN) policies to new environments is expensive because every additional route and instruction requires an embodied demonstration. Yet standard observation-to-action training uses only a small fraction of the information already contained in each trajectory. In particular, future obs...

Khang Nguyen, Hoang Pham Quang Nguyen, Ha Phuong Nguyen et al. · 0 citations
Preprint Sep 2026

EgoTSR++: Egocentric Spatiotemporal Reasoning for Task Progress Understanding

The results establish goal-conditioned state comparison as an explicit formulation of egocentric spatiotemporal reasoning for task-progress understanding as part of a unified framework for diagnosing and improving order-robust task-progress understanding.

Xiao-Da Yang, Can Wang, Yu-Xiang Liu et al. · 0 citations
Preprint Aug 2026

From Semantic Grounding to Decision Optimization: A Unified Framework for Long-Horizon UAV Vision-Language Navigation

This work presents an instruction-grounded semantic enhancement module that injects object-level semantics and relative spatial cues into the current observation state, and develops a relevance-aware dynamic temporal aggregation strategy that reweights the full history buffer while converting a few high-relevance frame...

Ze-Yuan Ma, Jiaxin Chen, Di Huang · 1 citation
Preprint Aug 2026

ARIES-Mission2: A Zero-Shot Vision-Language-Action Framework for Fast Large-Scale Aerial Mission Generation

This work proposes ARIES-Mission2, a zero-shot Vision-Language-Action (VLA) framework that decouples visual-semantic perception from physical route optimization in low-altitude Unmanned Aerial Vehicle (UAV) mission generation and indicates that the TSP module maintains lower growth in computation time as the number of...

Jun-Hao Wei, Yan-Xiao Li, Hao-Chen Li et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.