Skip to content
Preprint

NaviScale: Generating Large-Scale Semantic Map Datasets for Object Navigation

Sep 2026 · 0 citations · 35 references
Computer Science

TL;DR

NaviScale is proposed for semantic-map-based object navigation (ObjectNav), whose predictor can be trained on pairs of partial and complete semantic maps without reconstructing a complete 3D environment for every training sample.

Abstract

Embodied navigation requires spatial representations that generalize across unseen environments, yet collecting large amounts of annotated data from real 3D environments is difficult. We propose NaviScale for semantic-map-based object navigation (ObjectNav), whose predictor can be trained on pairs of partial and complete semantic maps without reconstructing a complete 3D environment for every training sample. The framework generates large-scale semantic map training data by composing floorplans of real homes with room-level semantic and obstacle maps extracted from MP3D and HM3DSem. NaviScale increases data diversity in two ways: inter-room scaling increases floorplan-level structural diversity, while intra-room scaling fills each fixed floorplan with different combinations of room maps matched by room category. Visibility through Ray Casting (VisRC) converts the composed maps into partial observations that account for field of view, sensing range, and occlusion. The resulting dataset contains 192,000 semantic maps generated from 24,000 floorplans associated with 12,794 properties. With 300k training iterations and the training and inference settings described in this paper, the system reaches 64.3% SR and 34.8% SPL on HM3D, together with 43.1% SR and 16.8% SPL on MP3D, without changing the prediction architecture. Additional experiments evaluate the quality of the composed maps, the effects of semantic-segmentation errors, and deployment on a physical robot.

View source

Similar papers

2026

TS-MapLoc: Large-Scale Indoor Object-Level Localization With Topological-Semantic Maps

Large-scale indoor mapping and positioning with vision sensors is fundamental to a wide range of applications, such as robotic navigation and augmented reality. However, the rapidly increasing number of detectable objects and the expanded spatial coverage jointly introduce matching ambiguity and high computational cost...

Cui-Yun Fang, Fan Wang, Ye-Dong Jiang et al. · 0 citations
Preprint Sep 2026

VoxelFix: Post-Hoc Semantic Correction of Completed 3D Voxel Maps

Semantic 3D maps are increasingly constructed automatically for aerial robotics by integrating learned semantic predictions into 3D representations. While this avoids costly manual 3D annotation, errors in the perception and mapping pipeline can persist in the resulting map, reducing its reliability for downstream auto...

Sunesh Praveen Raja Sundarasami, Taehyoung Kim, Johannes Scherer et al. · 0 citations
Conference Open access Sep 2026

Interactive Open-Set Semantic Mapping with a 3D Scene Graph Backend

A modular mapping architecture is demonstrated that establishes 3D Semantic Scene Graphs (3DSSGs) as its foundational back-end, enabling the dense representation of extensive environments containing thousands of unique object instances and supporting open-vocabulary queries via CLIP features without requiring any addit...

Felix Igelbrink, Lennart Niecksch, Martin Günther et al. · 0 citations
Preprint Aug 2026

LifelongCrossNav: Persistent 3D Semantic Memory for Cross-Floor Multi-Object Navigation

Experimental results show that LifelongCrossNav consistently outperforms a representative planar persistent semantic-map baseline on HM3D-MFMON, demonstrating that persistent 3D semantic memory and cross-floor traversability modeling effectively support sequential multi-object navigation in multi-floor environments.

Zehui Li, Zihao Sun, Jia-Wei Xu et al. · 0 citations
Preprint Sep 2026

Multi-Scale Semantic Mapping in Urban Environments via Observation Calibration and Policy Dependence Regularization

Semantic mapping is fundamental to embodied navigation, yet existing methods are developed for indoor environments, where objects exhibit relatively limited scale variation and are observed from a restricted range of viewpoints. Urban environments pose substantially greater challenges: agents must map objects ranging f...

Run-Ling Long, Jun-Hao Feng, Jia Wan · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.