This paper proposes a novel iterative refinement framework based on a video diffusion model to improve the completeness and consistency of dynamic 4D scenes, and substantially outperforms existing baselines.
Abstract
This paper addresses the challenges of dynamic scene synthesis from sparse-view videos. Existing methods employ geometric priors, adaptive optimization, or density-control strategies to improve 4D Gaussian modeling under sparse observations. However, they cannot fundamentally resolve the ill-posed problem caused by insufficient observations and missing scene information. Moreover, sparse-view 4D Gaussian Splatting (4DGS) often suffers from poor geometric initialization: with only a few input views, COLMAP typically reconstructs sparse and incomplete point clouds, leaving large scene regions without sufficient Gaussian support and making them difficult to recover through subsequent optimization. To address these limitations, we propose a novel iterative refinement framework based on a video diffusion model to improve the completeness and consistency of dynamic 4D scenes. Specifically, we first estimate multi-view depth maps and fuse them into dense point clouds to provide more complete geometric initialization for a dynamic 4DGS representation. We then employ a pretrained video restoration model to refine sequences rendered along novel camera trajectories at different time steps. The restored sequences serve as pseudo-supervision to regularize and iteratively refine the 4DGS representation. Experiments on a widely used benchmark dataset demonstrate that our method substantially outperforms existing baselines, achieving nearly a 2 dB PSNR improvement over the previous best-performing method.
This work proposes a novel 3D-aware video restoration framework designed to enhance the quality of sparse 3DGS reconstruction and introduces a camera-conditioned geometric prior that guides the network toward geometrically grounded restoration that remains coherent across viewpoints.
Xinhui Liu, Can Wang, Wei Jiang et al.· 1 citation
Novel view synthesis from sparse inputs remains challenging for 3D Gaussian Splatting (3DGS) due to ambiguous geometry, cross-view inconsistency, and missing details in under-constrained regions, resulting in degraded reconstruction and unstable rendering. To tackle these issues, we propose D$^{3}$GS, a Depth-DINO-Diff...
Yun-Qi Gao, Zhan-Feng Liao, Han-Zhang Tu et al.· 0 citations
This work introduces a geometry-aware cost volume that injects relative pose distance, view-dependent ray angle, and spatial validity masks into dense depth matching, enabling the network to jointly reason about photometric consistency, triangulation reliability, and visibility.
An alternating optimization framework that uses a pre-trained image diffusion model to generate geometrically consistent pseudo-views for additional 3DGS supervision and introduces Generative Active Pseudo-view Selection (GAPS) to balance reconstruction informativeness and generative reliability when choosing target vi...
Hong-Fei Zhu, Hao-Chen Deng, Si-Tao Zhang et al.· 0 citations
This work presents FixAnything, a single model for fixing a wide range of rendering artifacts by repurposing a pretrained video generative model, leveraging its implicit multi-view priors with only minimal modification and lightweight finetuning.
Khiem Vuong, D. Ramanan, Srinivasa G. Narasimhan· 1 citation
A novel framework designed to achieve high-fidelity dynamic neural scene reconstruction from highly sparse viewpoints by integrating geometry-aware depth priors and robust multi-view correspondence constraints is proposed, offering a scalable solution for dynamic scene capture.
Camila Wilson, D. J. Ross· Journal of innovative resear...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.