Skip to content

Self-Supervised Learning of Semantically Consistent Visual Features via a Simple 3D Prior

· 0 citations · 48 references

TL;DR

This work introduces a new self-supervised approach that allows us to learn features for topologically complex object categories using a simple prior, and makes use of contrastive learning and distribution matching at the global dataset-level to learn the coarse shape and appearance of a category.

View source

Similar papers

Aug 2026

VSLG-net: visual-spatial latent graph network for two-view correspondence learning

The Visual-Spatial Latent Graph Network (VSLG-Net), a parameter-compact transformer-based framework with dual-branch for local and global context perception in attention mechanisms, is proposed, which achieves competitive performance on the outdoor YFCC100M benchmark and remains competitive on the indoor SUN3D benchmar...

Wei Lv, Han-Lin Guo, Zhi Shen et al. · 0 citations
Preprint Sep 2026

Back to the Feature: Zero-Shot 6DoF Pose Estimation via Dense Local Features

B2TFPose is presented, a training-free zero-shot method for 6DoF pose estimation of unseen objects from RGB images, establishing state-of-the-art performance among training-free RGB methods and outperforming trained counterparts including GigaPose and GenFlow, at competitive inference speed.

Ali Rafiaei, Michael A. Greenspan · 0 citations
Preprint Sep 2026

LEGAU: Learning Semantic Gaussian Priors for Scalable Category-level Pose Estimation

Category-level 6D pose estimation from a single RGB-D observation is inherently under-constrained, since partial visible geometry must be interpreted together with a canonical object structure before a stable pose can be determined. We present LEGAU, a unified framework that jointly predicts NOCS correspondence, object...

Hong-Li Xu, Zhao-Wei Lu, Jun-Wen Huang et al. · 0 citations
Preprint Aug 2026

Emergent 3D Instance Segmentation from Self-Supervised Point Transformers

This work investigates whether a frozen, self-supervised point transformer already contains the structural information required to isolate object instances without any handcrafted geometric prior, and develops a training-free segmenter that groups points via connected components on a key-similarity graph, using neither...

Ted Lentsch, Santiago Montiel-Mar'in, Holger Caesar et al. · 0 citations
Preprint Aug 2026

DDMS: Discriminative Distillation of Multi-view Foundational Features into Single-view Models

Foundational visual features such as DINO have played a critical role across modern computer vision, and have recently become key components in multi-view feed-forward geometry estimators. In this work, we demonstrate that by re-distilling these multi-view models---their internal knowledge of 3D geometry---into a singl...

Jeong-gi Kwak, Sho Kagami, Yuki Ono et al. · 0 citations
Nov 2026

SparKLoc: Sparse Key-Gaussian Matching for 3D Gaussian-Based Visual Localization

3D Gaussian Splatting (3DGS) has emerged as an effective scene representation for visual localization, but most existing methods rely on embedding keypoint descriptors into Gaussian primitives, leading to high memory overhead and requiring joint optimization of geometry and features. We propose SparKLoc, a visual local...

Gyeong Chan Kim, Youngseok Jang, Jeong-Dae Heo et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.