Skip to content

SUFLECA: Scaling Up Feature Learning for CAD-to-image Alignment

Jul 2026 · arXiv.org · Vol abs/2607.15058 · 0 citations · 65 references
Computer Science

TL;DR

SUFLECA (Scaling Up Feature LEarning for CAD-to-image Alignment), a weakly supervised framework for zero-shot CAD alignment with two key contributions, and proposes a geometrically consistent matching algorithm that establishes reliable CAD-to-image correspondences.

Abstract

CAD-to-image alignment aims to estimate an object's 9D pose (rotation, translation, and anisotropic scale) from a single RGB image, with applications in robotics and augmented reality. Recent zero-shot methods use vision foundation models to match image regions to CAD models; yet their correspondences are typically appearance-driven or unreliable under occlusion or synthetic-to-real domain shift. To address these limitations, we introduce SUFLECA (Scaling Up Feature LEarning for CAD-to-image Alignment), a weakly supervised framework for zero-shot CAD alignment with two key contributions. First, SUFLECA scales up geometry-grounded feature learning from pretrained visual representations through Normalized Object Coordinates (NOCs) supervision on images spanning up to 12 real and synthetic datasets, learning compact geometry-aware features that generalize across domains. Second, we propose a geometrically consistent matching algorithm that establishes reliable CAD-to-image correspondences. Together, these contributions enable accurate, sub-second alignment per object instance without iterative pose refinement. On ScanNet25k, SUFLECA achieves 32.8%/42.6% category/instance accuracy, outperforming the strongest zero-shot baseline by 9.7/12.5 percentage points with a smaller computational footprint, and for the first time on this benchmark, even surpassing existing pose-supervised methods. Code is available at: https://github.com/snt-arg/SUFLECA

View source

Similar papers

Preprint Sep 2026

Back to the Feature: Zero-Shot 6DoF Pose Estimation via Dense Local Features

B2TFPose is presented, a training-free zero-shot method for 6DoF pose estimation of unseen objects from RGB images, establishing state-of-the-art performance among training-free RGB methods and outperforming trained counterparts including GigaPose and GenFlow, at competitive inference speed.

Ali Rafiaei, Michael A. Greenspan · 0 citations
Preprint Sep 2026

VGM-VS: Rethinking Visual Geometry Model for High-Precision Visual Servoing

We present VGM-VS, a visual servoing method built on a pretrained feed-forward visual geometry model. Given the current view and a reference image captured at the target configuration, we estimate the relative camera pose with a visual geometry model and apply it iteratively as the pose increment of a closed-loop pose-...

Yi-Min Pan, Sen-Miao Wang, You Zhou et al. · 0 citations
Conference 2026

Two-stage Monocular 6D Pose Estimation for Small Cubic Objects

This paper studies monocular 6D pose estimation of small cubic objects from a single RGB image and proposes a two-stage manipulation- oriented framework, which achieves the strongest overall balance in ADD-S, translation accuracy, rotation stability, and task-oriented usability metrics.

Xinmiao Du · 0 citations
Sep 2026

Unified feature Gaussian representation for robust camera relocalization

LoFG is presented, a localization-oriented Feature Gaussian representation that unifies both stages within a single Gaussian scene and improves the robustness of sparse initialization and the accuracy of dense refinement, demonstrating potential for localization applications in AR, robotics, and visual navigation syste...

Zhen-Dong Xiao, Zi-Ling Wen, Jun Yin et al. · 0 citations
Preprint Sep 2026

Revisiting Multi-View Stereo: A Sequence-to-Sequence Formulation

Computing accurate geometry from multi-view images is a fundamental problem in computer vision. Recent feed-forward (FF) models jointly estimate 3D geometry and camera parameters, but they typically suffer from geometry distortion caused by reconstruction ambiguity, even when ground-truth camera parameters are supplied...

Ao-Xiang Fan, Corentin Dumery, Nicolas Talabot et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.