Skip to content
Preprint

Guiding Image-to-3D Generation with Test-Time Partial Observations

Sep 2026 · 0 citations · 30 references
Computer Science

TL;DR

This work introduces a training-free framework for incorporating partial geometric observations into pretrained image-to-3D generative models without retraining or finetuning, and demonstrates that pretrained image-to-3D models can effectively integrate partial geometric observations through explicit test-time guidance.

Abstract

Image-to-3D models can generate visually compelling 3D assets from a single RGB image, but their geometry is often only loosely constrained by the available observations, limiting their use in applications that require geometric fidelity. In many real-world settings, however, partial geometric observations of the object may be available at test time. We introduce a training-free framework for incorporating such evidence into pretrained image-to-3D generative models without retraining or finetuning. To do this, we guide generation using a ray-consistent observation likelihood defined over the model's occupancy representation, combining surface occupancy and free-space evidence. Applied to SAM 3D and its multi-view extension, our approach substantially improves geometric fidelity across different levels of observability, as well as visual quality. Our results demonstrate that pretrained image-to-3D models can effectively integrate partial geometric observations through explicit test-time guidance, complementing their learned generative priors without modifying the underlying model.

View source

Similar papers

Preprint Aug 2026

SpatialCrafter: Single Image World Modeling with Generative 3D Proxies

SpatialCrafter is presented, a novel two-stage framework that addresses explorable image-to-scene generation issues by introducing a global 3D proxy for high-fidelity image-to-scene generation and appearance refinement and introduces Parallel Geometry Injection and Proxy-Aware Corruption training strategies.

Chuan Fang, Lingteng Qiu, Yixun Liang et al. · 1 citation
Preprint Oct 2026

Less Context, Better Geometry: Masked Geometric Encoder for Robust 3D Foundation Models

Recent progress in 3D foundation models has enabled rapid 3D reconstruction and camera calibration by leveraging learned 3D priors from vast amount of spatial data. However, the all-to-all global attention design leads to quadratic complexity and limits long-sequence inference; unconstrained cross-view interactions als...

Zhi-Min Shao, Xi-Jun Liu, Zhao-Liang Zhang et al. · 0 citations
Preprint Aug 2026

Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models

Self-Geometry, a plug-and-play test-time adaptation pipeline that directly imposes explicit multi-view geometric constraints using 2D pixel correspondences as pseudo ground-truth, achieves consistent improvements in both pose and geometry estimation across VFMs.

Seokhyun Youn, Dahyeon Kye, Sung-Ho Bae et al. · 0 citations
Preprint Aug 2026

ReconSplat: Generalizable 3D Scene Reconstruction Beyond Observed Views

We introduce ReconSplat, a feed-forward model for 3D scene reconstruction that aims to address the longstanding trade-off between plausible view generation for unobserved regions and geometric consistency, providing both geometrically aligned novel views and sharp depth estimates. Our approach builds on 3D Gaussian spl...

Giuseppe Stracquadanio, Kevin Raj, Julia Grabinski et al. · 0 citations
Preprint Aug 2026

WilLaGS: Latent-Conditional 3D Appearance Fields for Robust Gaussian Splatting In-the-Wild

A generative appearance model where a $\beta$-VAE learns a structured and continuous manifold of global appearance is introduced, Conditioned on the latent code, a 3D neural appearance field is constructed that generates dynamic Tri-Plane features to encode spatially-varying local illumination effects.

Yu Bai, Qian-Qiu Tan, Li-Long Chen et al. · 1 citation
Conference Aug 2026

Evaluating Diffusion Models for Single-Image 3D Gaussian Splatting Scene Reconstruction

Sparse or limited observations often cause 3D Gaussian Splatting (3DGS) scenes to exhibit holes, missing content, and unstable appearance in novel views. Repairing these regions requires generating synthetic frames that remain consistent across viewpoints, a challenge closely related to cross-view drift in generative m...

Max Eskandari, Xing-Bang Tang, Matthew J. Kyan · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.