Skip to content
Preprint

Unordered Landmark Visual Navigation

Aug 2026 · 0 citations · 84 references
Computer Science

TL;DR

Unordered Landmark Visual Navigation (ULVN) is proposed, a unified RGB-only framework free from temporal and odometric priors that constructs a robust 2D topological map directly from unstructured images via calibrated geometric verification and maximum spanning forest refinement.

Abstract

Image-goal navigation is a fundamental capability for embodied AI, yet its practical deployment is strained by strong prior assumptions. Existing methods predominantly rely on temporally ordered video streams or auxiliary sensors (e.g., depth, LiDAR) to maintain spatial consistency. These sequential and multimodal dependencies severely restrict scalability, especially when deploying robots using crowd-sourced or pre-recorded unordered image collections. When temporal priors are removed, current methods struggle with severe perceptual aliasing, noisy associations, and catastrophic mapping failures. To address this underexplored challenge, we propose Unordered Landmark Visual Navigation (ULVN), a unified RGB-only framework free from temporal and odometric priors. ULVN systematically mitigates error accumulation by integrating mapping, localization, and planning. Specifically, it constructs a robust 2D topological map directly from unstructured images via calibrated geometric verification and maximum spanning forest refinement. For closed-loop execution, ULVN abandons sequential heuristics, utilizing a graph-based belief propagation filter with entropy-adaptive fusion for global localization and dynamic subgoal planning. Extensive experiments in simulation and real-world deployments demonstrate that ULVN significantly outperforms state-of-the-art methods.

View source

Similar papers

Preprint Aug 2026

OccPlanner: Goal-Aware Occupancy-Conditioned Diffusion Planner for PixelGoal Navigation

PixelGoal navigation specifies targets directly in the agent's camera view, providing a natural interface between high-level visual reasoning and low-level navigation. Depth can lift a visible target pixel into a metric PointGoal, but this estimate becomes unreliable under occlusion or sensor noise. Moreover, a PointGoal alone does not encode traversability or feasible paths around obstacles. We present OccPlanner, a goal-aware occupancy-conditioned diffusion planner that learns complementary egocentric goal and planning-oriented 3D representations through metric target and occupancy prediction, respectively. These representations condition a diffusion trajectory module to generate target-directed, obstacle-aware trajectories. For scalable geometric supervision, we introduce L3ROcc, which converts monocular RGB navigation videos into aligned 3D occupancy and trajectory annotations. We train OccPlanner on L3ROcc-processed InternData-N1 and evaluate it in closed-loop simulation across four unseen InternScenes categories and two goal-distance ranges. Across all eight settings, OccPlanner substantially outperforms existing open-source PixelGoal approaches and achieves competitive performance against PointGoal planners with direct metric-goal inputs.

Binling Huang, Nianjin Ye, Xi Yang et al. · 0 citations
Preprint Sep 2026

Multi-Task Visual Perception Network with LLM Conditioning for Autonomous Navigation

Long-term navigation for service robots faces crit- ical challenges like the accumulation of odometry drift and sensor error, which progressively degrade 2D maps and renders traditional path planning algorithms (e.g., A*, RRT*, DiPPer, ViT-A*) ineffective over time. To address this, we propose a user-friendly, interactive framework that eliminates the reliance on globally consistent maps. Our approach integrates visual perception with Large Language Models (LLM) to interpret user commands via text or voice. Instead of relying on a drift- prone global map, the system generates a sequential action plan based on local visual cues and egocentric geometric instructions. These action plans are executed sequentially, allowing the robot to navigate known and unknown environments safely. By reset- ting localization relative to immediate targets, our framework effectively works with a minimum accumulation drift strategy, ensuring accurate, efficient, and collision-free navigation without the maintenance overhead of traditional mapping. Experiments on real-world and simulated data have shown significant improve- ments over other methods. Our source code is publicly accessible at https://github.com/PraveenSingh24/VL-Navigation.

Praveen Kumar, K. Guruprasad, Tushar Sandhan · 0 citations
Jul 2026

VidMap: Exploiting Temporal Structure for Video-Based Structure-from-Motion

This work introduces a system that combines the strong sequential constraints of SLAM with the flexibility and global optimization of offline SfM, enabling the metric reconstruction of arbitrary, long, uncalibrated videos.

Zador Pataki, Paul-Edouard Sarlin, Marc Pollefeys · 1 citation · ⚡1
Open access Aug 2026

EgoNav: Bridging Learned Waypoints and Geometry-Aware Local Control for Robust Indoor Navigation

Image-goal navigation using lightweight topological maps is a practical paradigm for indoor robot deployment: the map requires only geotagged images, and localization relies on visual matching rather than precise pose estimation. However, learned waypoint predictors can produce targets that violate geometric constraints or deviate from the global path. Executing these waypoints safely further requires a local planner capable of collision avoidance, yet existing systems either lack one or rely on fixed parameters that cannot adapt to confined spaces. To address these limitations while retaining the navigational intuition of the learned predictor, we present EgoNav, a hierarchical system that implements this idea by generating candidates from semantically segmented traversable regions and scoring them alongside the learned waypoint for geometric safety, directional coherence, and fidelity to the learned prior. An adaptive local path planner then executes the refined waypoint with parameters modulated based on the refinement outcome. Experiments in Habitat-sim and on a physical humanoid robot show that EgoNav consistently outperforms contemporary baselines in both success rate and path efficiency.

Jing Wang, Shiqi Zhao, Hai-Rong Qu et al. · 0 citations
Oct 2026

ForexNav: Foresight Exploratory Navigation in Complex and Unknown Indoor Environments

Autonomous navigation in unknown, complex indoor environments remains challenging due to limited sensing range and severe partial observability. Conventional methods rely on local maps without foresight, causing dead-ends and long detours, while local goal selection based on Euclidean distance or frontier coverage fails to balance efficiency with directionality. To address these challenges, we propose ForexNav, a foresight-enabled exploratory navigation framework. To handle structural ambiguity in unseen regions, we introduce Foresight Hypothesis Fusion (FHF), which maintains multiple WGAN-based map predictions and reweights them via temporal evidence accumulation. A Traversability-aware A* search then quantifies predictive traversability on the fused map, enabling a multi-objective planner to synthesize path feasibility, kinodynamic conformity, monotonic-progress consistency, and geometric distance for optimal intermediate goal selection and dynamically consistent trajectory generation. Experiments in four simulated indoor scenes of up to 3,300 m<inline-formula><tex-math notation="LaTeX">$^{2}$</tex-math></inline-formula> demonstrate navigation success while reducing total travel time by 25.0% and improving average velocity by 13.3% over the strongest baseline, with path ratio improvements of 22.2% on average in large-scale environments (<inline-formula><tex-math notation="LaTeX">$\geq$</tex-math></inline-formula>2,000 m<inline-formula><tex-math notation="LaTeX">$^{2}$</tex-math></inline-formula>). Real-world deployment on a quadruped robot supports practical feasibility, and extension to a fixed-altitude micro-UAV further suggests preliminary cross-platform transferability.

Hong-Yu Song, Yun-Fang Ren, Ji-Gui Miao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Dual-Layer Semantic-Spatial Belief Mapping for Aerial Object Goal Navigation

Aerial Object Goal Navigation (ObjectNav) requires an unmanned aerial vehicle (UAV) to locate a described target in an unknown outdoor environment using onboard visual observations. Vision-language models (VLMs) can interpret open-ended target descriptions and visual observations, but their frame-level outputs are often noisy, sparse, and spatially transient. We propose AeroBelief, a dual-layer semantic-spatial belief mapping framework that transforms transient VLM observations into persistent spatial guidance. It separates broad contextual plausibility from target-specific evidence: an intuition layer accumulates scene-level semantic cues for exploration, while an evidence layer preserves qualified target-specific observations for approach and confirmation. Evidence-gated fusion combines the two layers into spatial belief hotspots. We further introduce object-conditioned visual reasoning with conservative evidence qualification to improve observation reliability before spatial accumulation. In parallel, egocentric regional guidance converts quadtree coverage into UAV-centered, yaw-aligned directional proposals and stabilizes them through temporal commitment. Its regional scoring is independent of semantic belief values, maintaining exploration pressure and reducing repeated low-gain search. Experiments on the UAV-ON benchmark show that AeroBelief achieves the best reported overall SR, OSR, and SPL among the compared methods, reaching 21.61%, 35.57%, and 10.62, respectively. These results support the effectiveness of persistent semantic-spatial belief, conservative evidence qualification, and temporally stable regional guidance for aerial ObjectNav.

Jian-Qiang Xiao, Xiang Deng, Yue-Xuan Sun et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.