Jul 2026· 2026 23rd International Conference on Ubiquitous Robots (UR)· pp. 183-188· 0 citations· 26 references
Computer Science
Abstract
3D scene graphs provide structured environmental representations that enable robots to perform language-grounded tasks such as navigation and manipulation. A key challenge in constructing 3D scene graphs is preserving consistent object identity. During robot navigation, newly observed 3D segments need to be associated with previously accumulated instances. Existing methods perform this association using geometric overlap and feature similarity, but these cues alone progressively fragment or merge object instances falsely. To address this challenge, we present GORI, an image-guided selective 3D object re-association framework that combines 2D multi-object tracking with selective 3D re-association. GORI employs 2D temporal tracking as the primary association mechanism and performs 3D re-association selectively to account for tracking discontinuity. To mitigate erroneous merges of spatially adjacent objects, GORI enforces a co-detection constraint that prevents merging 3D instances observed as distinct detections within the single image. The resulting 3D scene graph provides consistent object instances that serve as a grounding space for language-conditioned task planning. We evaluate GORI on HM3DSem dataset over existing 3D scene graph baselines, and assess its performance on real-world indoor scans. We demonstrate improved panoptic quality, F1 score, and average precision on HM3DSem; showing that improved object consistency supports more robust downstream task planning.
NavPatch is presented, an object level correction layer that assigns ADD, REMOVE, or EXTEND to navigation relevant object categories through periodic scene understanding with a vision-language model.
Shi-Jie Sun, Xing-Yu Tao, Hao Wang et al.· 0 citations
The framework, GauScoreMap, performs hierarchical scoring: a training-free global stage produces a CLIP relevance field over Gaussians and prunes the scene to a small set of candidate regions, and a local stage performs ray-image cross-attention within those regions and triangulates the 6D camera pose from the top-scor...
Yijie Deng, Shuaihang Yuan, Geeta Chandra Raju Bethala et al.· IEEE Robotics and Automation...· 4 citations· ⚡1
SceneBench is introduced, a benchmark of 966 photorealistic 3D scenes reconstructed with Gaussian Splatting and densely annotated with hierarchical semantics spanning scenes, rooms, functional areas, object groups, and individual objects that provides a realistic testbed for developing and evaluating models capable of...
Anubhav Khanal, Prabigya Acharya, Roshni Poudel et al.· 0 citations
This work introduces PLANET, an end-to-end multi-object tracker designed to move beyond the image plane, and lifts existing 2D tracking datasets into 3D by embedding reconstructed 3D scene geometry into the features and positional encodings used during query formation.
Orcun Cetintas, Guillem Brasó, Tim Meinhardt et al.· 0 citations
Open-vocabulary 3D maps enable robots to reason about previously unknown environments using natural language. However, existing systems typically segment every incoming image, associate detections with persistent 3D segments, and frequently perform costly Vision-Language (VL) inference. We present TRACKGRAPH, an online...
Peder Borge Hellesylt, Albert Gassol Puigjaner, Kostas Alexis et al.· 0 citations
Recent 3D large multimodal models (3D-LMMs) rely on a visual bottleneck to compress complex 3D scene evidence into a limited number of visual tokens compatible with large language models (LLMs). Current visual bottlenecks, however, often passively compress heterogeneous 3D evidence into a homogeneous object-centric tok...
Xiang-Qi Li, Li-Bo Huang, Jia-Rui Zhao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.