NavVerse is introduced, a physics-enabled benchmark for indoor-to-outdoor embodied navigation, and zero-shot experiments with RL, VLA, and modular baselines show that current agents remain far from solving cross-context navigation.
Abstract
Robots deployed in delivery, campus, and emergency-response settings often need to navigate from buildings to streets within a single continuous episode. Existing benchmarks usually evaluate indoor and outdoor navigation separately, and many abstract away robot execution, leaving exit finding, boundary traversal, adaptation, and kinodynamic failures underexplored. We introduce NavVerse, a physics-enabled benchmark for indoor-to-outdoor embodied navigation. NavVerse contains 100 indoor scenes, 50 urban outdoor scenes, and 50 indoor-to-outdoor scenes, and 10,000 episodes spanning Object Navigation, Vision-and-Language Navigation, and Place Navigation tasks, where agents search for semantic points of interest such as restaurants or banks. Agents are evaluated through executable robot interfaces using task-success, path-efficiency, and safety metrics. Zero-shot experiments with RL, VLA, and modular baselines show that current agents remain far from solving cross-context navigation: end-to-end VLAs obtain the highest zero-shot success, while the modular method provides the strongest safety profile. PlaceNav further reveals a clear drop from outdoor to indoor-to-outdoor scenes, indicating that adaptation remains major bottleneck.
This paper proposes LSTP-Nav, a lightweight, decentralized navigation framework built on LSTP-Net that maps stacked 2D LiDAR observations, goal information, and velocity feedback directly to action and introduces an HS reward to provide smooth, heading-aware safety feedback, and develops PhysReplay-SimLab to improve tr...
Xingrong Diao, Zhi-Qiang Sun, Jian-Wei Peng et al.· IEEE Transactions on Automat...· 0 citations
Long-duration outdoor coverage with autonomous platforms remains challenging beyond classical planning: deployments face localization drift in open spaces, obstacles in cluttered sites, controller feasibility in turn-heavy maneuvers, and persistent autonomy with energy management. We propose a unified ROS 2 architectur...
L. Gargani, Matteo Frosi, Matteo Matteucci· 0 citations
Vision-and-language navigation (VLN) has advanced rapidly in static indoor environments, but robots operating in human-populated spaces must ground language while responding to moving pedestrians and social-safety constraints. We present DPed-VLN, a Habitat 3.0 benchmark for dynamic-pedestrian VLN that couples 33,093 n...
Hao-Jie Dai, Xiang-Yi Wang, Liu-Yi Wang et al.· 0 citations
Mapless navigation often removes global maps while retaining localization-derived goal vectors or bearings. We study a stricter setting in which a mobile robot observes only local LiDAR, scalar goal range, and short histories of executed actions; neither pose nor goal direction is provided to the policy. We introduce A...
Autonomous navigation in unknown, complex indoor environments remains challenging due to limited sensing range and severe partial observability. Conventional methods rely on local maps without foresight, causing dead-ends and long detours, while local goal selection based on Euclidean distance or frontier coverage fail...
Hong-Yu Song, Yun-Fang Ren, Ji-Gui Miao et al.· IEEE Robotics and Automation...· 0 citations
The results support the feasibility of using geometrically screened, temporally maintained doorway cues to modify frontier scheduling; direct traversable-area coverage, detector accuracy, runtime, and real-robot performance remain outside the present evidence.