A mask-aware tri-modal framework that improves the quality of superpoint representations by retrieving a scene-level structural context from a pretrained PointSAM encoder to enhance object-centric evidence and predicting a soft mask weight to suppress unreliable superpoints.
For embodied agents to navigate and reason indoor spaces, they need object-level 3D representations that stay consistent over time as new frames arrive from a monocular camera. Current online 3D instance segmentation methods either depend on posed RGB-D input with ground-truth depth or couple tightly to the internal re...
M. Mohrat, E. Derevyanka, I. Obrubov et al.· Journal of Instrument Engine...· 0 citations
OccBEV-Oracle, a lightweight completion neck inserted between the BEV backbone and dense head of a CenterPoint-style detector, and a direct measurement of where the neck edits the BEV map, show that the aligned GT-ROM mask itself supplies a strong localization prior.
Jun Wang, Quanxin Zheng, Jian-Ping Yu· IEEE Access· 0 citations
This work investigates whether a frozen, self-supervised point transformer already contains the structural information required to isolate object instances without any handcrafted geometric prior, and develops a training-free segmenter that groups points via connected components on a key-similarity graph, using neither...
Ted Lentsch, Santiago Montiel-Mar'in, Holger Caesar et al.· 0 citations
OVR-GS (Open-Vocabulary Removal in Gaussian Splatting), an instruction-driven object-removal framework for pre-trained 3D Gaussian Splatting (3DGS) scenes, demonstrates the effectiveness of localized Gaussian optimization for instruction-driven cleanup of reconstructed environments before visual inspection, presentatio...
Yongpeng Ding, Feng Ouyang, Jiawei Fan et al.· Italian National Conference...· 0 citations
This work enhances the existing iterative object-basesd visual localization approach with an additional semantic feature derived from a pretrained semantic segmentation model and conducts a systematic baseline study of contemporary feature matching techniques on such cross-domain query-reference image pairs.
Yasmin Loeper, Markus Gerke, P. Fanta-Jende· The International Archives o...· 0 citations
Results indicate that incorporating textual semantic priors can effectively enhance high-level semantic representations of point clouds, providing a feasible solution for indoor 3D scene understand.
Jinyu Tan, Juntao Yang, Yutao Zhang et al.· The International Archives o...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.