Skip to content

Mask-aware tri-modal learning for indoor 3D object detection

Jul 2026 · The Visual Computer · Vol 42 · 0 citations · 41 references
Computer Science

TL;DR

A mask-aware tri-modal framework that improves the quality of superpoint representations by retrieving a scene-level structural context from a pretrained PointSAM encoder to enhance object-centric evidence and predicting a soft mask weight to suppress unreliable superpoints.

View source

Similar papers

Open access Jul 2026

SAM3R: Object-Centric 3D Mapping via Foundation-Model-Guided Data Association in Changing Scenes

For embodied agents to navigate and reason indoor spaces, they need object-level 3D representations that stay consistent over time as new frames arrive from a monocular camera. Current online 3D instance segmentation methods either depend on posed RGB-D input with ground-truth depth or couple tightly to the internal re...

M. Mohrat, E. Derevyanka, I. Obrubov et al. · 0 citations
Open access 2026

Decoupling Mask Quality From Completion Design: A Diagnostic Framework for Occlusion-Aware LiDAR 3-D Object Detection

OccBEV-Oracle, a lightweight completion neck inserted between the BEV backbone and dense head of a CenterPoint-style detector, and a direct measurement of where the neck edits the BEV map, show that the aligned GT-ROM mask itself supplies a strong localization prior.

Jun Wang, Quanxin Zheng, Jian-Ping Yu · 0 citations
Preprint Aug 2026

Emergent 3D Instance Segmentation from Self-Supervised Point Transformers

This work investigates whether a frozen, self-supervised point transformer already contains the structural information required to isolate object instances without any handcrafted geometric prior, and develops a training-free segmenter that groups points via connected components on a key-similarity graph, using neither...

Ted Lentsch, Santiago Montiel-Mar'in, Holger Caesar et al. · 0 citations
Open access Aug 2026

OVR-GS: Open-Vocabulary 3D Object Removal via Semantic Gaussian Selection and Local Diffusion-Guided Completion

OVR-GS (Open-Vocabulary Removal in Gaussian Splatting), an instruction-driven object-removal framework for pre-trained 3D Gaussian Splatting (3DGS) scenes, demonstrates the effectiveness of localized Gaussian optimization for instruction-driven cleanup of reconstructed environments before visual inspection, presentatio...

Yongpeng Ding, Feng Ouyang, Jiawei Fan et al. · 0 citations
Open access Aug 2026

Using textureless, low-detailed 3D city models for visual localization

This work enhances the existing iterative object-basesd visual localization approach with an additional semantic feature derived from a pretrained semantic segmentation model and conducts a systematic baseline study of contemporary feature matching techniques on such cross-domain query-reference image pairs.

Yasmin Loeper, Markus Gerke, P. Fanta-Jende · 0 citations
Open access Aug 2026

Text-Guided Semantic Segmentation Method for Indoor 3D Point Clouds

Results indicate that incorporating textual semantic priors can effectively enhance high-level semantic representations of point clouds, providing a feasible solution for indoor 3D scene understand.

Jinyu Tan, Juntao Yang, Yutao Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.