This work introduces a new self-supervised approach that allows us to learn features for topologically complex object categories using a simple prior, and makes use of contrastive learning and distribution matching at the global dataset-level to learn the coarse shape and appearance of a category.
The Visual-Spatial Latent Graph Network (VSLG-Net), a parameter-compact transformer-based framework with dual-branch for local and global context perception in attention mechanisms, is proposed, which achieves competitive performance on the outdoor YFCC100M benchmark and remains competitive on the indoor SUN3D benchmar...
Wei Lv, Han-Lin Guo, Zhi Shen et al.· Signal, Image and Video Proc...· 0 citations
B2TFPose is presented, a training-free zero-shot method for 6DoF pose estimation of unseen objects from RGB images, establishing state-of-the-art performance among training-free RGB methods and outperforming trained counterparts including GigaPose and GenFlow, at competitive inference speed.
Category-level 6D pose estimation from a single RGB-D observation is inherently under-constrained, since partial visible geometry must be interpreted together with a canonical object structure before a stable pose can be determined. We present LEGAU, a unified framework that jointly predicts NOCS correspondence, object...
Hong-Li Xu, Zhao-Wei Lu, Jun-Wen Huang et al.· 0 citations
This work investigates whether a frozen, self-supervised point transformer already contains the structural information required to isolate object instances without any handcrafted geometric prior, and develops a training-free segmenter that groups points via connected components on a key-similarity graph, using neither...
Ted Lentsch, Santiago Montiel-Mar'in, Holger Caesar et al.· 0 citations
Foundational visual features such as DINO have played a critical role across modern computer vision, and have recently become key components in multi-view feed-forward geometry estimators. In this work, we demonstrate that by re-distilling these multi-view models---their internal knowledge of 3D geometry---into a singl...
Jeong-gi Kwak, Sho Kagami, Yuki Ono et al.· 0 citations
3D Gaussian Splatting (3DGS) has emerged as an effective scene representation for visual localization, but most existing methods rely on embedding keypoint descriptors into Gaussian primitives, leading to high memory overhead and requiring joint optimization of geometry and features. We propose SparKLoc, a visual local...
Gyeong Chan Kim, Youngseok Jang, Jeong-Dae Heo et al.· IEEE Robotics and Automation...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.