Jul 2026· International Conference on Signal Processing and Communications· pp. 1-5· 0 citations· 26 references
Abstract
Standard 3D Gaussian Splatting (3DGS) pipelines for Novel View Synthesis (NVS) are bottlenecked by Structurefrom-Motion (SfM) initialization. In casual, sparse-view scenarios, feature matching breaks down, causing the entire reconstruction process to fail. We replace this brittle dependency with a COLMAP-free, feed-forward initializer powered by a Visual Geometry Grounded Transformer (VGGT). By leveraging VGGT, our pipeline jointly estimates camera parameters and dense scene geometry across all views in a single pass. A Bridge Module then robustly normalizes the scene scale and conditions initial Gaussian opacity on geometric confidence to discourage floater artifacts during densification. Our framework reduces the initialization phase from minutes (full-scene SfM) to seconds and achieves $\mathbf{1 0 0} \boldsymbol{\%}$ initialization success from as few as three unposed images (a regime where COLMAP succeeds on only 1 of 7 Mip-NeRF 360 scenes). Project page: https://github.com/yuvanrajkrishna/VGGT-Sparse-3DGS.
Across RealEstate10K, DL3DV, Tanks-and-Temples, and Mip-NeRF 360, SplatGuide achieves state-of-the-art pose-free novel view synthesis, and surpasses the ground-truth-pose baseline.
Ye-Jun Zhang, Zi-Han Wang, Xue-Si Ji et al.· 0 citations
A semantic-guided 3D Gaussian splatting (3DGS) framework tailored to sparse-view industrial reconstruction was introduced, enabling robust reconstruction from limited viewpoints and offers a practical geometric foundation for automated inspection and remote equipment monitoring.
Boyang Li, Tian-Han Gao, Zuan Gu et al.· Visual Computing for Industr...· 0 citations
InfoLoD introduces a Fisher-guided self-distillation scheme that uses the Fisher Information Matrix to select geometrically valid, information-rich pseudo viewpoints, enabling LoD training directly from a pre-trained 3DGS model without any original images.
Zhenyu Xia, Pengcheng Han, Lin Chen et al.· IEEE Transactions on Visuali...· 0 citations
FlexSplat matches or approaches posed state-of-the-art reconstructors while requiring neither camera poses nor ground-truth depth, and matches the best perceptual (LPIPS) quality among the compared methods on GSO.
Amir Sabbaghziarani, Han-Ting Ye, Maria Gorlatova et al.· 0 citations
SARG-GS is proposed, a geometry-driven 3DGS framework tailored for sparse-view scenarios, comprising a Semantic Augmented Epipolar Fusion (SAEF) module and a Residual Guided Reprojection Compensation (RRC) module, which achieves superior structural completeness and rendering fidelity with as few as three input views.
Huan Zhou, Huizhi Zhu, Jiongming Qin et al.· The Visual Computer· 0 citations
3D Gaussian Splatting (3DGS) achieves high-quality, real-time novel view synthesis, but the resulting assets have baked-in illumination and cannot be easily relit. Inverse rendering methods optimize simplified reflectance and illumination models for each scene, limiting efficiency and relighting quality. Recent generative approaches leverage large diffusion models for realistic lighting edits, but applying them to 3DGS typically requires an additional per-scene optimization stage to bake the edited appearance into the representation. We present LightBridge, a feed-forward generative framework for controllable relighting of complete 3DGS assets in a single pass. To enable feed-forward training, we construct a large-scale Multi-Illumination Relighting Dataset with paired source and target observations of the same scenes. Latent Bridge Relighting Diffusion models relighting as source-to-target transport in latent space, enabling one-step extraction of 2D visual tokens without iterative diffusion sampling. A Gaussian Propagation Transformer uses a point transformer with sparse image-to-point self-attention followed by point-to-image cross-attention to efficiently propagate these cues across the complete 3DGS, while avoiding full attention over all image and Gaussian tokens. Experiments validate these designs, demonstrating competitive relighting quality and efficient single-pass prediction of complete relit 3DGS assets without scene-specific optimization. The code and dataset will be made publicly available upon acceptance.
Heng Cao, Pan-Hao Cheng, Huang-Sheng Du et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.