Skip to content
Preprint

PePESeg3D: Perception Prior Enhances Multi-Scale Segmentation for 3D Gaussian Splatting

Sep 2026 · 0 citations · 49 references
Computer Science

TL;DR

PePESeg3D achieves state-of-the-art performance in both multi-scale segmentation and scene reconstruction, highlighting the importance of integrating perception priors into both geometry optimization and feature learning for accurate multi-scale 3D segmentation.

Abstract

Recent advancements in 3D Gaussian Splatting (3DGS) have extended its capabilities to multi-scale segmentation. Existing methods reconstruct a scene with Gaussian primitives and learn multi-scale segmentation features separately, which leaves the geometry unaware of semantic structure and the feature learning dependent on incomplete mask supervision. To address these limitations, we present PePESeg3D, a novel framework that injects perception priors into a multi-scale 3D Gaussian segmentation pipeline. To fully exploit perception priors, we integrate them not only into contrastive feature learning but also into the upstream geometry reconstruction. Specifically, PePE Reconstruction incorporates monocular depth and mask constraints to ensure semantically coherent object structures. Building on this aligned geometry, PePE Contrastive Learning leverages dense depth-color cues and view-consistent centroid supervision to compensate for the incompleteness of multi-scale masks obtained from a 2D foundation model. Extensive experiments on the SPIn-NeRF, LERF-Mask, and NVOS benchmarks demonstrate that PePESeg3D achieves state-of-the-art performance in both multi-scale segmentation and scene reconstruction, highlighting the importance of integrating perception priors into both geometry optimization and feature learning for accurate multi-scale 3D segmentation. Our code is available at https://github.com/BeCow5X5/PePESeg3D.

View source

Similar papers

Preprint Aug 2026

GaussianDS: Depth-supervised Semantic Gaussian Splatting for Scene Understanding

GaussianDS, a depth-supervised semantic 3DGS framework that treats semantic lifting as a supervision-alignment problem and jointly optimizes RGB appearance, rendered depth, and compact semantics from scratch, is proposed.

Yu-Fei Zhang, Chen-Lu Zhan, Hong-Wei Wang · 0 citations
Preprint Aug 2026

GaussVid: Sparse-View Gaussian Splatting with 3D-Aware Video Diffusion Priors

This work proposes a novel 3D-aware video restoration framework designed to enhance the quality of sparse 3DGS reconstruction and introduces a camera-conditioned geometric prior that guides the network toward geometrically grounded restoration that remains coherent across viewpoints.

Xinhui Liu, Can Wang, Wei Jiang et al. · 1 citation
Open access Sep 2026

Semantic-Guided Adaptive Gaussian Segmentation

3D Gaussian Splatting (3DGS) enables real-time photorealistic scene reconstruction, yet its segmentation tasks suffer from two critical flaws: poor 3D consistency (e.g., blurred instance boundaries and unstable cross-view semantic association) and insufficient structural awareness near ambiguous object boundaries. To a...

Ying-Han Zhou, Fan Zhou · 0 citations
Preprint Sep 2026

D3GS: Depth, DINO, and RGB Diffusion Co-Guided 3D Gaussian Splatting for Sparse-View Reconstruction

Novel view synthesis from sparse inputs remains challenging for 3D Gaussian Splatting (3DGS) due to ambiguous geometry, cross-view inconsistency, and missing details in under-constrained regions, resulting in degraded reconstruction and unstable rendering. To tackle these issues, we propose D$^{3}$GS, a Depth-DINO-Diff...

Yun-Qi Gao, Zhan-Feng Liao, Han-Zhang Tu et al. · 0 citations
Sep 2026

XMask3D++: Cross-Modal Mask Reasoning for Open Vocabulary 3D Segmentation.

Existing methodologies in open vocabulary 3D semantic segmentation primarily concentrate on establishing a unified feature space encompassing 3D, 2D, and textual modalities. Nevertheless, traditional techniques such as global feature alignment or vision-language model distillation tend to impose only approximate corres...

Ziyi Wang, Yan-Bo Wang, Xu-Min Yu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.