Skip to content

Annotation-Free Furniture Codes: What They Encode, and How Far They Transfer

Jul 2026 · arXiv.org · Vol abs/2607.10461 · 0 citations · 27 references
Computer Science

TL;DR

This work asks whether a single self-supervised token derived from object geometry can replace both a categorical class label and a canonical-pose convention, and study such tokens directly as a representation, decoupled from any synthesizer.

Abstract

Layout-based 3D scene synthesizers place each object using two human-annotated channels: a categorical class label and a canonical-pose convention. We ask whether a single self-supervised token derived from object geometry can replace both, and study such tokens directly as a representation, decoupled from any synthesizer. A Finite Scalar Quantization (FSQ) point-cloud autoencoder is chamfer-trained on placed 3D-FUTURE furniture with no labels or pose annotations. Diagnostic probes recover fine-category (62.6 +/- 0.5%), super-category (85.6 +/- 1.3%), and yaw (52.7 +/- 0.5 deg) from the codes alone. Swapping the chamfer target from the rotated to the un-rotated point cloud collapses the yaw signal while raising class recovery, showing the codes'rotation content can be set by the training objective. Scaling across asset libraries needs codes that transfer; on an unseen dataset (ShapeNet), alignment is category-dependent: box-like furniture transfers, organically-shaped furniture does not, and a target-blind augmentation partly closes the gap.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

GoDeep: Annotation-Free Open-Vocabulary 3D Scene Understanding via Language-Space Lifting

Open vocabulary 3D semantic segmentation methods typically lift CLIP features into 3D. This embeds points in a joint vision-language space known to behave like a bag-of-words on compositional tasks. Furthermore, even annotation free variants often require a large 3D training corpus and a dedicated 3D encoder per domain...

T. Betsas, A. Doulamis, Andreas Georgopoulos · 0 citations
Preprint Aug 2026

PoseAdapter: Dual-Stream 2.5D Controllable Image Generation for Complex Multi-Object Scenes

PoseAdapter, a lightweight framework for high-fidelity 2.5D controllable image generation, and a Context-Aware Dual-Stream Representation, to resolve the generative trade-off between strict instance isolation and global coherence.

Yu-Feng Chi, Hui-Min Ma, Fan Gao et al. · 0 citations
Preprint Aug 2026

SR-JEPA: Learning Predictive Latent State in 3D Scenes

Joint-embedding predictive architectures learn by predicting latent representations of missing observations, yet many masked JEPAs are evaluated primarily through the encoders they produce. We ask what a trained predictive pathway itself infers when an entire entity is absent from a native 3D scene. We introduce SR-JEP...

Zihan Zhou, Qi-Fu Wen, Xi Zeng · 0 citations
Jul 2026

The Lift Spectrum: How Measurement-to-Space Adaptivity Shapes Robustness in Image-Free Single-Pixel Sensing

It is shown that removing image reconstruction relocates the central design problem to the lift: how 1D measurements become a 2D task representation, and measurement-to-space adaptivity organizes both the clean-to-noisy operating envelope and the failure a system encounters.

Yu-Yuan Han, Jing-Wei Li, Xiaoxia Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.