ForesightNav is presented, a training-free universal zero-shot navigation framework that adaptively combines context-aware semantic value mapping, graph-based target verification, and TSP-based frontier ordering and results further support the effectiveness of adaptive switching, context-aware semantic mapping, and TSP-based frontier ordering.
Abstract
Universal zero-shot goal-oriented navigation requires an agent to locate object categories, target instances, or text-described goals in unseen environments without task-specific policy training. Existing methods usually rely on either broad semantic exploration or fine-grained target verification, but they can become inefficient when semantic cues are sparse, ambiguous, or insufficient for reliable graph matching. This paper presents ForesightNav, a training-free universal zero-shot navigation framework that adaptively combines context-aware semantic value mapping, graph-based target verification, and TSP-based frontier ordering. Guided by semantic-score dispersion and graph-matching confidence, the planner switches among geometric exploration, semantic exploration, partial-match exploration, and target verification, balancing broad goal-directed search with fine-grained target confirmation. Experiments on MP3D, HM3D, and RoboTHOR across Object Navigation (ObjectNav), Instance Navigation (InstanceNav), and Text Navigation (TextNav) show favorable performance across different goal modalities. Compared with UniGoal, ForesightNav improves ObjectNav performance on three evaluated benchmarks, with the largest gains on RoboTHOR ObjectNav: 18.6 percentage points in Success Rate (SR) and 11.4 in Success weighted by Path Length (SPL). On HM3D, it achieves 65.3% SR on InstanceNav and 29.6% SR on TextNav, suggesting that the same training-free planning process remains applicable to instance- and language-specified goals. Foundation-model analysis and ablation results further support the effectiveness of adaptive switching, context-aware semantic mapping, and TSP-based frontier ordering. Code is available at https://anonymous.4open.science/r/ForesightNav
PixelGoal navigation specifies targets directly in the agent's camera view, providing a natural interface between high-level visual reasoning and low-level navigation. Depth can lift a visible target pixel into a metric PointGoal, but this estimate becomes unreliable under occlusion or sensor noise. Moreover, a PointGo...
Binling Huang, Nian-Jin Ye, Xi Yang et al.· 0 citations
Semantically-Guided Exploration (SGE) is introduced, a modular exploration framework for ground vehicles that integrates pixel-level semantic segmentation into sampling-based waypoint selection and receding-horizon route optimization and introduces mechanisms to address real-world navigation uncertainty.
This work proposes an object--path graph that unifies open-vocabulary semantic reasoning with topological navigation, and introduces a navigation strategy that combines global path planning with local inter-node execution through lightweight node localization and semantic visual servoing, enabling navigation directly o...
Lin-Wei Zheng, Dao-Jie Peng, Bing-Tao Wang et al.· 0 citations
EgoPathBench, a dataset and five-task benchmark for first-person waypoint decision-making, and Fine-tuning Qwen 3.5 4B on the released training split improves all four reported evaluations across three external spatial benchmarks, with gains of 1.4--9.6 points.
Results indicate that the proposed perception-to-control framework can support language-grounded target approach manoeuvres of ASV under the complex port environments and demonstrate the importance of semantic grounding, harbour-aware filtering, and semantic verification for reliable language-grounded ASV navigation.
This work presents an instruction-grounded semantic enhancement module that injects object-level semantics and relative spatial cues into the current observation state, and develops a relevance-aware dynamic temporal aggregation strategy that reweights the full history buffer while converting a few high-relevance frame...
Ze-Yuan Ma, Jiaxin Chen, Di Huang· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.