Results show that the latent space is central to geometry-consistent generation and can serve as a shared interface between perception and generation.
Abstract
We present a compact geometry-native latent space as a shared foundation for perception and generation. Visual generators can produce photorealistic frames without preserving a consistent 3D scene. We argue that this is not only a modeling problem but also a representation problem: generators typically evolve appearance-centric latents, while perception models recover geometry in a semantically rich space that encodes cross-view structure. Rather than adding geometry as another output, we reparameterize a geometry foundation model's features into a compact latent space for generation. We realize this shift with the geometry-native autoencoder (GAE), whose latent is jointly decodable to appearance, depth, cameras, and point maps. With this state, a standard conditional flow supports diverse generation tasks. In controlled comparisons that hold the generator and training protocol fixed, replacing the latent with GAE improves both visual quality and independently measured 3D coherence: FVD falls by $12.7\%$ and $23.1\%$ on RealEstate10K and DL3DV, and camera-trajectory error is halved on RealEstate10K. Together, these results show that the latent space is central to geometry-consistent generation and can serve as a shared interface between perception and generation.
A generative appearance model where a $\beta$-VAE learns a structured and continuous manifold of global appearance is introduced, Conditioned on the latent code, a 3D neural appearance field is constructed that generates dynamic Tri-Plane features to encode spatially-varying local illumination effects.
Yu Bai, Qian-Qiu Tan, Li-Long Chen et al.· 0 citations
Novel view synthesis from sparse images must reconcile faithful reconstruction of observed regions with plausible completion of unseen content, while maintaining world consistency across viewpoints. Existing geometry-based methods preserve observed scene structure but often struggle to complete unseen regions, whereas...
Ke-Rui Ren, Tao Lu, Lin-Ning Xu et al.· 0 citations
SpatialCrafter is presented, a novel two-stage framework that addresses explorable image-to-scene generation issues by introducing a global 3D proxy for high-fidelity image-to-scene generation and appearance refinement and introduces Parallel Geometry Injection and Proxy-Aware Corruption training strategies.
Chuan Fang, Lingteng Qiu, Yixun Liang et al.· 1 citation
High-fidelity image-to-3D generation requires a 3D representation that captures both geometry and appearance. However, preserving fine detail across the physically based rendering (PBR) modalities needed for relighting remains challenging. To address this, we propose Luce, a 3D representation that unifies geometry and...
M. Singh, Michele Stoppa, Alvise Memo et al.· 0 citations
GaussianWAM is proposed, a training-time representation-enhancement framework that organizes geometric and semantic supervision through a 3D Gaussian field and improves performance on standard LIBERO and shows positive transfer trends on RoboTwin and real-world manipulation.
Zi-Jian Zhang, Yu-Qing Jiang, Wei-Tao Zhou et al.· 1 citation
Transferring deformation between characters with different geometry and topology is challenging because conventional rigs encode behaviour through character-specific structures and correspondences. We present LatentReRig, an experimental framework that investigates whether pose-associated changes can instead be represe...
D. Dolci, Fabrizio Poggioni, Carlo Melchiorri· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.