Skip to content
Conference

GORI: Image-Guided Selective 3D Object Re-Association for 3D Scene Graphs and Task Planning

Jul 2026 · 2026 23rd International Conference on Ubiquitous Robots (UR) · pp. 183-188 · 0 citations · 26 references
Computer Science

Abstract

3D scene graphs provide structured environmental representations that enable robots to perform language-grounded tasks such as navigation and manipulation. A key challenge in constructing 3D scene graphs is preserving consistent object identity. During robot navigation, newly observed 3D segments need to be associated with previously accumulated instances. Existing methods perform this association using geometric overlap and feature similarity, but these cues alone progressively fragment or merge object instances falsely. To address this challenge, we present GORI, an image-guided selective 3D object re-association framework that combines 2D multi-object tracking with selective 3D re-association. GORI employs 2D temporal tracking as the primary association mechanism and performs 3D re-association selectively to account for tracking discontinuity. To mitigate erroneous merges of spatially adjacent objects, GORI enforces a co-detection constraint that prevents merging 3D instances observed as distinct detections within the single image. The resulting 3D scene graph provides consistent object instances that serve as a grounding space for language-conditioned task planning. We evaluate GORI on HM3DSem dataset over existing 3D scene graph baselines, and assess its performance on real-world indoor scans. We demonstrate improved panoptic quality, F1 score, and average precision on HM3DSem; showing that improved object consistency supports more robust downstream task planning.

View source

Similar papers

Open access Jun 2025

Hierarchical Scoring With 3D Gaussian Splatting for Instance Image-Goal Navigation

The framework, GauScoreMap, performs hierarchical scoring: a training-free global stage produces a CLIP relevance field over Gaussians and prunes the scene to a small set of candidate regions, and a local stage performs ray-image cross-attention within those regions and triangulates the 6D camera pose from the top-scor...

Yijie Deng, Shuaihang Yuan, Geeta Chandra Raju Bethala et al. · 4 citations · ⚡1
#artificial intelligence Preprint Sep 2026

SceneBench: A Hierarchical Benchmark for Vision-Language Understanding of 3D Scenes

SceneBench is introduced, a benchmark of 966 photorealistic 3D scenes reconstructed with Gaussian Splatting and densely annotated with hierarchical semantics spanning scenes, rooms, functional areas, object groups, and individual objects that provides a realistic testbed for developing and evaluating models capable of...

Anubhav Khanal, Prabigya Acharya, Roshni Poudel et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond the Image Plane: World-Grounded Queries for Multi-Object Tracking

This work introduces PLANET, an end-to-end multi-object tracker designed to move beyond the image plane, and lifts existing 2D tracking datasets into 3D by embedding reconstructed 3D scene geometry into the features and positional encodings used during query formation.

Orcun Cetintas, Guillem Brasó, Tim Meinhardt et al. · 0 citations
Preprint Sep 2026

TRACKGRAPH: Online Open-Vocabulary 3D Scene Graphs via Image-Space Tracking

Open-vocabulary 3D maps enable robots to reason about previously unknown environments using natural language. However, existing systems typically segment every incoming image, associate detections with persistent 3D segments, and frequently perform costly Vision-Language (VL) inference. We present TRACKGRAPH, an online...

Peder Borge Hellesylt, Albert Gassol Puigjaner, Kostas Alexis et al. · 0 citations
Preprint Sep 2026

SceneScaffold: Active Scene-State Construction for Unified 3D Scene Understanding

Recent 3D large multimodal models (3D-LMMs) rely on a visual bottleneck to compress complex 3D scene evidence into a limited number of visual tokens compatible with large language models (LLMs). Current visual bottlenecks, however, often passively compress heterogeneous 3D evidence into a homogeneous object-centric tok...

Xiang-Qi Li, Li-Bo Huang, Jia-Rui Zhao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.