Skip to content
Preprint

GraphWrit3R: End-to-End 3D Scene Graph Writing

Sep 2026 · 0 citations · 83 references
Computer Science

TL;DR

This work presents GraphWrit3R, a simple end-to-end method that takes a 3D point cloud, Gaussian Splats, or a combination of both as input, and directly outputs a complete scene graph as a structured JSON script, achieving state-of-the-art performance on object class, predicate, and triplet recall on the 3DSSG benchmark.

Abstract

3D scene graphs provide a structured representation of complex environments by encoding objects, their semantic attributes, and the spatial and functional relationships between them. Current approaches for 3D scene graph generation suffer from several fundamental limitations. They rely on complex multi-stage pipelines with explicit intermediate representations, making systems fragile and prone to error propagation. They assume access to ground-truth object annotations during inference, which deviates from real-world scenarios. They depend on proprietary models, hindering open-source deployment, or incur prohibitively slow inference. We present GraphWrit3R, a simple end-to-end method that takes a 3D point cloud, Gaussian Splats, or a combination of both as input, and directly outputs a complete scene graph as a structured JSON script. The graph lists all objects, their semantic attributes, and the relationships between them, while avoiding all of the above mentioned limitations. The choice of multiple input modalities is purely for versatility, allowing a single set of weights to handle diverse scenarios. Point cloud inputs are encoded via Sonata and Gaussian Splat inputs via Chorus, with both modalities projected onto a shared voxel grid and fused through a novel per-voxel contrastive alignment loss before being decoded by a large language model. As a natural consequence of the LLM, GraphWrit3R also supports open-vocabulary querying. On the 3DSSG benchmark, our method achieves state-of-the-art performance on object class, predicate, and triplet recall, outperforming methods that rely on ground-truth object annotations during inference. We further provide qualitative results and analyze different input modality configurations, contrastive loss formulations, and token fusion strategies.

View source

Similar papers

Preprint Sep 2026

Yggdrasil: a Layer-First 3D Scene Graph for Real-Time Querying

Robotic agents use 3D scene graphs (3DSG) to perform tasks ranging from scene understanding to scene interaction. Although an extensive body of work addresses scene graph generation, little attention has been paid to optimizing the graph for consumption, which leaves state-of-the-art perception pipelines to work around...

Arshia Akhavan, Ermanno Bartoli, Afnan Algharbi et al. · 0 citations
Conference Open access Sep 2026

Interactive Open-Set Semantic Mapping with a 3D Scene Graph Backend

A modular mapping architecture is demonstrated that establishes 3D Semantic Scene Graphs (3DSSGs) as its foundational back-end, enabling the dense representation of extensive environments containing thousands of unique object instances and supporting open-vocabulary queries via CLIP features without requiring any addit...

Felix Igelbrink, Lennart Niecksch, Martin Günther et al. · 0 citations
Preprint Sep 2026

Hierarchical Aggregation of Semantic Uncertainty in 3D Scene Graphs

Open-vocabulary 3D Scene Graphs (3DSGs) ground each object node in a vision-language embedding, yet they record every entry as equally certain, so a robot querying the map cannot tell which of its entries are unreliable. Estimators of semantic uncertainty could supply that distinction, but they require repeated samplin...

Carlos Roberto Cueto Zumaya, Iacopo Catalano, W. Bessa et al. · 0 citations
Review Open access Aug 2026

Semantic 3D Gaussian Splatting: A State-of-the-Art Review

A unified multi-axis taxonomy is introduced that enables us to classify the available methods in 3D Gaussian splatting methods in terms of five complementary categories: semantic vocabulary space, representation form, functional role, knowledge source, and query mechanism.

J. Flotyński · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.