This work introduces a geometry-aware cost volume that injects relative pose distance, view-dependent ray angle, and spatial validity masks into dense depth matching, enabling the network to jointly reason about photometric consistency, triangulation reliability, and visibility.
Abstract
Generalizable 3D Gaussian Splatting enables efficient sparse-view novel view synthesis, but accurate Gaussian center estimation remains challenging. Epipolar attention and cost volume methods exhibit complementary limitations in non-Lambertian regions, appearance-ambiguous areas, and scenes with large perspective changes. These limitations lead to feature mismatches, depth estimation errors, and geometric distortions. To address these limitations, we propose GeoSplat, a feed-forward Generalizable 3D Gaussian Splatting framework with geometry-aware priors. Specifically, we introduce a geometry-aware cost volume that injects relative pose distance, view-dependent ray angle, and spatial validity masks into dense depth matching, enabling the network to jointly reason about photometric consistency, triangulation reliability, and visibility. Furthermore, we design a ray-guided iterative refinement module, in which full-resolution 3D ray direction and depth confidence priors jointly guide recurrent residual updates to progressively refine coarse depth predictions in a continuous space. Extensive experiments on RealEstate10K and ACID demonstrate that GeoSplat achieves competitive reconstruction quality with a compact parameter count, while presenting an accuracy efficiency trade-off.
This paper proposes a novel iterative refinement framework based on a video diffusion model to improve the completeness and consistency of dynamic 4D scenes, and substantially outperforms existing baselines.
Hai-Tao Huang, Sheng-Hao Zhao, Bo-Yuan Tian et al.· 0 citations
Novel view synthesis from sparse inputs remains challenging for 3D Gaussian Splatting (3DGS) due to ambiguous geometry, cross-view inconsistency, and missing details in under-constrained regions, resulting in degraded reconstruction and unstable rendering. To tackle these issues, we propose D$^{3}$GS, a Depth-DINO-Diff...
Yun-Qi Gao, Zhan-Feng Liao, Han-Zhang Tu et al.· 0 citations
This work proposes GSPotential, a framework that quantifies view-space supervision imbalance using a Camera Potential Field, and uses the potential field to guide reconstruction from two complementary aspects.
Zeyuan An, Yang Xiao, Zhiying Leng et al.· 0 citations
This work proposes a novel 3D-aware video restoration framework designed to enhance the quality of sparse 3DGS reconstruction and introduces a camera-conditioned geometric prior that guides the network toward geometrically grounded restoration that remains coherent across viewpoints.
Xinhui Liu, Can Wang, Wei Jiang et al.· 1 citation
Remote sensing novel view synthesis under sparse observations remains challenging due to insufficient geometric constraints and limited cross-view supervision. Existing Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) methods are prone to overfitting and face challenges of depth ambiguities, missing cross...
Experiments show that CoMVS-GS remains competitive on object-level reconstruction and improves geometric accuracy and mesh compactness in outdoor scenes while maintaining high rendering quality.
Shi-Han Chen, Junjing Zhang, Q. Yan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.