Skip to content

MonoVoc: Decoupling Geometry and Semantics for Lightweight Monocular Open-Vocabulary 3D Gaussians

Jul 2026 · arXiv.org · Vol abs/2607.28300 · 0 citations · 35 references
Computer Science

TL;DR

This work presents a novel, training-free pipeline that fundamentally reimagines this paradigm by explicitly decoupling 3D geometric reconstruction from semantic integration, and delivers an order-of-magnitude reduction in memory usage compared to SOTA baselines.

Abstract

Open vocabulary 3D scene understanding is essential for next-generation interactive systems, empowering users to intuitively query and navigate reconstructed environments using natural language. However, current 3D Gaussian frameworks are often bottlenecked by restrictive multiview capture requirements, costly scene-specific optimization, and the massive memory overhead of storing dense language features. We present a novel, training-free pipeline that fundamentally reimagines this paradigm by explicitly decoupling 3D geometric reconstruction from semantic integration. Given a standard monocular video sequence as input, our method efficiently outputs a compact, highly interpretable, and fully searchable object-level semantic Gaussian map. Rather than entangling heavy language embeddings within the mapping loop, we extract geometry independently and ground semantics through a lightweight, modular post-processing framework. Extensive evaluations on the Replica dataset demonstrate that this decoupled architecture preserves strong rendering fidelity and competitive segmentation accuracy. Crucially, by replacing dense per-Gaussian storage with modular, object-level semantic embeddings, our approach delivers an order-of-magnitude reduction in memory usage compared to SOTA baselines. This provides a highly efficient, scalable, and practical solution for open-vocabulary 3D retrieval and question answering directly from everyday monocular video.

View source

Similar papers

Preprint Aug 2026

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding

Qwen-3D is introduced, a geometry-aware LMM that compresses visual information within the Qwen backbone using multi-view geometric cues, enabling efficient long-horizon visual reasoning over static scenes and incorporates a query-based segmentation decoder that grounds language directly in the underlying 3D scene repre...

Lucy Lin, Ayush Jain, Yifan Liu et al. · 3 citations
Preprint Aug 2026

Dynamic-Robust Photometric-Semantic Reconstruction for Open-Vocabulary 3D Scene Understanding

SPAR is proposed, a novel joint semantic-geometric encoding architecture that explicitly isolates transient dynamic noise prior to latent space aggregation that structurally couples motion estimation with multi-view visual and semantic learning and reveals a strong inter-task synergy between photometric scene reconstru...

Boyu Cai, Li Yang, Yan Xu et al. · 0 citations
#machine learning Preprint Sep 2026

SparseTalk - Sparsifying 3D Gaussian Language Fields for Efficient 3D Visual Question Answering

3D Gaussian language fields provide an explicit, spatially grounded representation for 3D visual question answering (VQA), but their dense semantic features can require tens of thousands of embeddings per scene, resulting in substantial storage, memory, and inference costs. We investigate how much of this representatio...

Davit Soselia, J. Jaja, Amitabh Varshney · 0 citations
Preprint Aug 2026

GaussianDS: Depth-supervised Semantic Gaussian Splatting for Scene Understanding

GaussianDS, a depth-supervised semantic 3DGS framework that treats semantic lifting as a supervision-alignment problem and jointly optimizes RGB appearance, rendered depth, and compact semantics from scratch, is proposed.

Yu-Fei Zhang, Chen-Lu Zhan, Hong-Wei Wang · 0 citations
Open access Sep 2026

Vision–language guided semantic-geometric transformer for memory-efficient 3D scene understanding

This work proposes a segmentation framework guided by Contrastive Language–Image Pre-training (CLIP) that enriches sparse 3D tokens with vision–language semantic priors and introduces a decoupled CLIP-induced semantic residual that forms semantic-geometric attention biases for local window attention.

Li-Cheng Liu, Yu Li, Fu-Yong Liu · 0 citations
Conference Open access Sep 2026

Interactive Open-Set Semantic Mapping with a 3D Scene Graph Backend

A modular mapping architecture is demonstrated that establishes 3D Semantic Scene Graphs (3DSSGs) as its foundational back-end, enabling the dense representation of extensive environments containing thousands of unique object instances and supporting open-vocabulary queries via CLIP features without requiring any addit...

Felix Igelbrink, Lennart Niecksch, Martin Günther et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.