Skip to content
Preprint

CanonNav: Disentangling Navigation Behavior from Camera Geometry in Cross-Platform Visual Navigation

Aug 2026 · 1 citation · 29 references
Computer Science

TL;DR

CanonNav is a visual navigation framework that disentangles navigation behavior from camera geometry and incorporates complementary planning supervision into learning from cross-platform demonstrations and introduces camera geometry canonicalization, which transforms visual observations and trajectories into a camera-consistent representation space.

Abstract

While visual navigation has advanced through imitation learning from cross-platform demonstrations, fully leveraging such data remains challenging. First, directly learning from image-trajectory pairs entangles navigation behavior with platform-dependent camera geometry. This hinders consistent learning by forcing the policy to implicitly infer camera geometry from visual observations, an inherently ill-posed problem. Second, imitation learning from demonstrated trajectories captures the expert's chosen motion but leaves the intermediate decisions underlying that motion implicit. To address these issues, we propose CanonNav, a visual navigation framework that disentangles navigation behavior from camera geometry and incorporates complementary planning supervision into learning from cross-platform demonstrations. CanonNav introduces camera geometry canonicalization, which transforms visual observations and trajectories into a camera-consistent representation space. Building on this representation, we derive safety and local-progress supervision using pseudo-labels from an offline traversability estimator. Safety supervision penalizes unsafe trajectories, while local-progress supervision guides where the robot should advance. Experiments across diverse camera configurations and environments show that, despite using only RGB at inference, CanonNav consistently outperforms RGB-based baselines and even surpasses RGB-D-based methods in challenging scenarios.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

DiffWAM: A Fast and Efficient Navigation World Action Model

Pretrained video foundation models encode rich semantic and spatiotemporal priors for embodied navigation, yet converting these priors into UAV motion typically requires expensive future-video synthesis and geometric reconstruction. We investigate whether the motion implicit in future visual prediction can instead be r...

Morui Zhu, Yu-Ze Wu, Xi-Jie Huang et al. · 0 citations
Preprint Sep 2026

ReVNM: Learning-Based Visual Navigation from a Remote Camera

Visual Navigation Models (VNMs) enable robots to navigate from egocentric visual observations without geometric localization and planning, but long-range navigation still requires pre-built maps. This paper presents the Remote Visual Navigation Model (ReVNM), which uses a single remote surveillance camera to serve as b...

Michikuni Eguchi, Kohei Honda, Masafumi Endo et al. · 0 citations
Preprint Oct 2026

Sensor-Layout-Agnostic Navigation via Geometric Observation Canonicalization

Existing visual navigation policies are inherently bound to fixed camera configurations, creating a fundamental barrier to zero-shot deployment across heterogeneous robot sensor layouts. To overcome this limitation, we present an embodiment-informed navigation policy capable of generalizing across diverse depth sensor...

Welf Rehberg, Kostas Alexis · 0 citations
Preprint Sep 2026

RoGSW4RLD: Feed-Forward 4D Gaussian Lifting for Robot World Model Rollouts

Action-conditioned video world models predict future robot interactions from multiple cameras, yet their outputs remain disparate video collections rather than a shared metric scene queryable across viewpoints and time. While existing 4D reconstruction methods offer a path to spatialize these predictions, independently...

Jin Hyun Kim, Min Young Kim, Soohwan Song et al. · 0 citations
Preprint Oct 2026

GeoScaffold: Learning Compact Geometric Latents via Reconstruction for Efficient Vision-Language Navigation

Recent vision-and-language navigation (VLN) systems increasingly adopt streaming Video-LLM policies that map egocentric RGB observations and instructions directly to low-level actions. Yet these policies inherit weak 3D geometric priors from 2D pretraining. Existing geometry-aware extensions charge a persistent inferen...

Yi-Xuan Jiang, Wen-Tong Li, An Liu et al. · 0 citations
Preprint Oct 2026

StageVLN: Spatial and Trajectory Auxiliary Guidance for Efficient Vision-Language Navigation

Vision-and-Language Navigation (VLN) policies increasingly benefit from strong semantic priors provided by large vision-language models (VLMs). However, standard action supervision does not explicitly encourage intermediate representations to preserve scene geometry, relative orientation, or global episode progress. In...

Anh Dao, Q. Phạm, Le Danh Vinh et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.