Skip to content

PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space

Jul 2026 · arXiv.org · Vol abs/2607.05373 · 2 citations
Computer Science

TL;DR

PixWorld is introduced, a single model that jointly addresses 3D reconstruction and generation that consistently outperforms prior latent-space generation methods and matches state-of-the-art reconstruction methods, demonstrating the superiority of a unified pixel-space approach.

Abstract

3D reconstruction and generation are commonly tackled by separate paradigms: pixel-based regression for reconstruction, and latent diffusion for generation. Recent works attempt to unify them in latent space, but with notable drawbacks: the diffusion objective is defined on latent features rather than the underlying 3D representation, and both branches suffer from information loss introduced by latent encoding, while requiring a pretrained Variational Autoencoder (VAE) or Representation Autoencoder (RAE). In this paper, we reformulate these two tasks under a unified pixel-space diffusion paradigm and introduce PixWorld, a single model that jointly addresses 3D reconstruction and generation. By supervising diffusion directly on rendered images, PixWorld removes the above limitations and aligns optimization with 3D scene fidelity. Beyond photometric and perceptual supervision that operates at the 2D image level and lacks 3D geometric awareness, we further introduce a geometry perception loss that aligns rendered views with their ground truth in the geometry-aware feature space of a pretrained 3D foundation model, providing 3D structural supervision. PixWorld consistently outperforms prior latent-space generation methods and matches state-of-the-art reconstruction methods, demonstrating the superiority of a unified pixel-space approach.

View source

Similar papers

Preprint Aug 2026

ReconSplat: Generalizable 3D Scene Reconstruction Beyond Observed Views

We introduce ReconSplat, a feed-forward model for 3D scene reconstruction that aims to address the longstanding trade-off between plausible view generation for unobserved regions and geometric consistency, providing both geometrically aligned novel views and sharp depth estimates. Our approach builds on 3D Gaussian spl...

Giuseppe Stracquadanio, Kevin Raj, Julia Grabinski et al. · 0 citations
#computer vision Preprint Aug 2026

Luce: Relightable Gaussians for 3D Asset Generation

High-fidelity image-to-3D generation requires a 3D representation that captures both geometry and appearance. To support relighting and integration into standard rendering pipelines, the representation should include physically based rendering (PBR) modalities such as albedo, metallic-roughness, and surface normals. We...

M. Singh, Michele Stoppa, Alvise Memo et al. · 0 citations
Preprint Aug 2026

SpatialCrafter: Single Image World Modeling with Generative 3D Proxies

SpatialCrafter is presented, a novel two-stage framework that addresses explorable image-to-scene generation issues by introducing a global 3D proxy for high-fidelity image-to-scene generation and appearance refinement and introduces Parallel Geometry Injection and Proxy-Aware Corruption training strategies.

Chuan Fang, Lingteng Qiu, Yixun Liang et al. · 1 citation
Open access Oct 2025

SaLon3R: Structure-Aware Long-Term Feedforward 3D Reconstruction from Unposed Images

This work proposes SaLon3R, a novel framework for Structure-aware, Long-term 3DGS Reconstruction that effectively prunes the redundant 3DGS and resolves artifacts in a single feed-forward pass, and introduces a 3D Point Transformer to overcome geometric inconsistencies caused by long-term accumulative errors.

Jiaxin Guo, Tongfan Guan, Wen-Zhen Dong et al. · 5 citations · ⚡1
Jul 2026

Axolotl3D: a Unified Framework for Faithful 3D Shape Completion

Axolotl3D is presented, a multi-modal and occlusion-aware 3D generation model that jointly conditions on images, visibility masks, camera parameters, and a partial point cloud that synthesizes diverse conditioning regimes from large-scale 3D data, enabling robust cross-modal reasoning.

A. Hu, Maria Shugrina · 1 citation
Preprint Aug 2026

Beyond Pixels: From Video Priors to 4D Worlds

Direct latent-to-4D generation is introduced and instantiate it as Latent-to-4D, which bypasses RGB by aligning a video latent with the token grid of a pretrained 4D decoder and refining it through frame-wise and global spatiotemporal attention.

Zihao Liu, Xi Shen, Zhen Zhou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.