Skip to content
Open access

DevGRU: Depth-Guided Visual Navigation Using a Collision-Aware Recurrent Model

Aug 2026 · IEEE Robotics and Automation Letters · Vol 11, pp. 12384-12391 · 0 citations · 31 references
Computer Science

TL;DR

The proposed DevGRU navigation system employs an action predictor that generates collision-aware future trajectories, enabling effective avoidance of immediate obstacles and has a relatively small number of trainable parameters, resulting in the fastest inference time among the baselines.

Abstract

Existing visual navigation models often aim to develop foundation models that can generalize robot navigation across diverse platforms. However, many of these models are prone to collisions when deployed in complex indoor environments, particularly in structured layouts and narrow passages. To address this problem, we propose a depth image- and point-goal-conditioned navigation system, DevGRU. The proposed system employs an action predictor (AP) that generates collision-aware future trajectories, enabling effective avoidance of immediate obstacles. In conjunction with a collision predictor, the AP further compensates for errors accumulated in the goal pose estimation and proactively mitigates future deviations. To evaluate our method, we conducted experiments across nine different scenes and three state-of-the-art approaches - ViNT, NoMaD, and NavDP - as well as four additional variants of ViNT and NoMaD. In terms of navigation performance, DevGRU significantly outperforms ViNT and NoMaD by a large margin. In addition, the proposed model has a relatively small number of trainable parameters, resulting in the fastest inference time among the baselines, particularly outperforming NavDP by 7× in model size and 17× in inference time.

Read PDF

Similar papers

Preprint Sep 2026

Multi-Task Visual Perception Network with LLM Conditioning for Autonomous Navigation

Long-term navigation for service robots faces crit- ical challenges like the accumulation of odometry drift and sensor error, which progressively degrade 2D maps and renders traditional path planning algorithms (e.g., A*, RRT*, DiPPer, ViT-A*) ineffective over time. To address this, we propose a user-friendly, interactive framework that eliminates the reliance on globally consistent maps. Our approach integrates visual perception with Large Language Models (LLM) to interpret user commands via text or voice. Instead of relying on a drift- prone global map, the system generates a sequential action plan based on local visual cues and egocentric geometric instructions. These action plans are executed sequentially, allowing the robot to navigate known and unknown environments safely. By reset- ting localization relative to immediate targets, our framework effectively works with a minimum accumulation drift strategy, ensuring accurate, efficient, and collision-free navigation without the maintenance overhead of traditional mapping. Experiments on real-world and simulated data have shown significant improve- ments over other methods. Our source code is publicly accessible at https://github.com/PraveenSingh24/VL-Navigation.

Praveen Kumar, K. Guruprasad, Tushar Sandhan · 0 citations
Preprint Aug 2026

OccPlanner: Goal-Aware Occupancy-Conditioned Diffusion Planner for PixelGoal Navigation

PixelGoal navigation specifies targets directly in the agent's camera view, providing a natural interface between high-level visual reasoning and low-level navigation. Depth can lift a visible target pixel into a metric PointGoal, but this estimate becomes unreliable under occlusion or sensor noise. Moreover, a PointGoal alone does not encode traversability or feasible paths around obstacles. We present OccPlanner, a goal-aware occupancy-conditioned diffusion planner that learns complementary egocentric goal and planning-oriented 3D representations through metric target and occupancy prediction, respectively. These representations condition a diffusion trajectory module to generate target-directed, obstacle-aware trajectories. For scalable geometric supervision, we introduce L3ROcc, which converts monocular RGB navigation videos into aligned 3D occupancy and trajectory annotations. We train OccPlanner on L3ROcc-processed InternData-N1 and evaluate it in closed-loop simulation across four unseen InternScenes categories and two goal-distance ranges. Across all eight settings, OccPlanner substantially outperforms existing open-source PixelGoal approaches and achieves competitive performance against PointGoal planners with direct metric-goal inputs.

Binling Huang, Nianjin Ye, Xi Yang et al. · 0 citations
Conference Open access Jul 2026

ZONDA: Zero-shot Object Navigation with Dynamic Avoidance in Multi-floor Environments

ZONDA, a zero-shot object navigation with dynamic avoidance framework, integrates three core components and can maintain robust navigation on the dynamic benchmark HM3D-DYNA compared to the existing baseline.

Shao-Min Liang, Xuan-Hong Liao, Shi-Yao Zhang · 0 citations
Preprint Aug 2026

Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation

This work proposes TAMP-Nav, a unified framework for efficient embodied navigation that dynamically triggers Chain-of-Thought and retains high-fidelity memory only at critical nodes, compressing redundant trajectories into lightweight Space-Time Indicators, thereby preserving critical historical information and enhancing spatio-temporal perception.

Hongyan Feng, Sunlai Chen, Xuanyu Liu et al. · 1 citation
Preprint Aug 2026

UniNav: A Unified World-Action Diffusion Model for Visual Navigation

Image-goal visual navigation is a fundamental capability for embodied agents. Existing navigation policies efficiently predict waypoint trajectories but lack visual foresight, while navigation world models can anticipate future observations but often require costly planning rollouts. We present UniNav, a unified world-action model that generates future visual observations and continuous waypoint trajectories through a single diffusion process. Given history frames and a goal image, UniNav jointly denoises visual and waypoint tokens within a single transformer, unifying future prediction and action generation in a shared framework. To improve spatial grounding, we incorporate geometry-aware camera tokens. We also train on both trajectory-labeled navigation data and video-only data, enabling the model to benefit from diverse videos without waypoint annotations. Based on this unified framework, we introduce two variants: UniNav-Full jointly predicts interpretable future observations and their corresponding trajectories, while UniNav-Fast removes future-image tokens at inference for efficient trajectory prediction. Experiments on navigation benchmarks show that UniNav outperforms the strongest baseline in ATE across all datasets. With one-step inference, UniNav-Fast achieves a latency of 0.1s without a substantial accuracy drop. Code will be released.

Changqing Zhou, Yueru Luo, Zeyu Jiang et al. · 0 citations
Jul 2026

VoLN: Vision-Only Long-Horizon Navigation - Paradigm, Benchmark, and Method

This work instantiates VoLN for aerial navigation through VoLN-UAV, a 7,210-episode benchmark that combines long-horizon goal-directed flight, continuous 3D motion, large viewpoint changes, and context-dependent beacon selection and reveals substantial remaining challenges in long-horizon evidence integration, cross-view goal matching, and closed-loop stability.

Jiabin Lou, Hao-Peng Wang, Yuan-Shuai Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.