Skip to content
Preprint

LEGO: Leveled Language Gaussian Splatting

Aug 2026 · 0 citations · 63 references
Computer Science

TL;DR

This work introduces LEGO for advanced open-vocabulary scene understanding, and establishes new state-of-the-art performance across both promptable and open-vocabulary 3D segmentation benchmarks, exhibiting advanced hierarchical scene decomposition and context-aware spatial reasoning.

Abstract

We introduce LEGO for advanced open-vocabulary scene understanding. Beyond basic concept recognition, its core innovation lies in capturing the intrinsic semantic hierarchies within the scene, such as the"flowerpot ->bouquet ->bud ->petal"lineage. While foundation models like SAM can identify multi-granular structures in 2D, their partitions are strictly perspective-bound and lack cross-view consensus. LEGO self-adaptively re-grades volatile multi-view SAM granularities into a unified, 3D-consistent hierarchy. This provides precise supervision for the structurally coherent, multi-level segmentation of 3D scenes. By grounding these segments with CLIP embeddings, LEGO recovers open-vocabulary semantic logic across hierarchical levels. Furthermore, by incorporating spatial relationships, we elevate these segments into level-wise language scene graphs, effectively empowering Large Language Models to perform complex, context-aware spatial reasoning and precise visual grounding. Experimental results demonstrate that LEGO establishes new state-of-the-art performance across both promptable and open-vocabulary 3D segmentation benchmarks, exhibiting advanced hierarchical scene decomposition and context-aware spatial reasoning.

View source

Similar papers

Preprint Aug 2026

Modeling Scientific Experiment Scenes: Dataset and Model

The Cross-Modal Dual-Path Generator (CM-DPG), a model for robust open-vocabulary SGG that enhances object-level semantic representations through joint visual-textual encoding and improves relational reasoning using complementary visual and geometric cues, is proposed.

Ming-Hao Zou, Qingtian Zeng, Shangkun Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

TDDN: Text-aligned Diffused DINO Network for Puzzle Understanding

This work fuse DINOv3 and CleanDIFT representations into a perception encoder and align it with RoBERTa-L, yielding a text-aligned model TDDN that preserves this perceptual advantage and leads on segmentation benchmarks among general-purpose contrastive encoders, including SigLIP.

Harsha Patnala, Debopriyo Banerjee, A. Munot et al. · 0 citations
Preprint Aug 2026

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding

Qwen-3D is introduced, a geometry-aware LMM that compresses visual information within the Qwen backbone using multi-view geometric cues, enabling efficient long-horizon visual reasoning over static scenes and incorporates a query-based segmentation decoder that grounds language directly in the underlying 3D scene repre...

Lucy Lin, Ayush Jain, Yifan Liu et al. · 3 citations
Jul 2026

MonoVoc: Decoupling Geometry and Semantics for Lightweight Monocular Open-Vocabulary 3D Gaussians

This work presents a novel, training-free pipeline that fundamentally reimagines this paradigm by explicitly decoupling 3D geometric reconstruction from semantic integration, and delivers an order-of-magnitude reduction in memory usage compared to SOTA baselines.

Pouya Ardekhani, Zahra Dehghanian, Morteza Abolghasemi et al. · 0 citations
Preprint Aug 2026

Hi-Token: Hierarchical Coordinate Tokenization for Generative Visual Grounding

The results show that structured coordinate generation provides an effective approach to generative visual grounding and Hi-Token encodes each coordinate with axis-specific tokens for the hundreds, tens, and ones digits, which adds coarse-to-fine structure and increases token reuse while retaining the existing VLM arch...

Xiuyuan Zhu, Ke Lu, Kun Dong et al. · 0 citations
Open access Aug 2026

RoFLIP: Robust and Fine-Grained Alignment for Vision-Language Compositional Reasoning

The Robust and Fine-grained training framework for CLIP-based vision-language models (RoFLIP) is proposed, enhancing both the robustness and granularity of vision-language alignment and underscore RoFLIP’s compositional reasoning and generalization abilities.

Yiwei Sun, Chuan-Bin Liu, Shancheng Fang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.