The key contribution is the application of Riemannian Flow Matching to 4D Gaussian Splatting parameters, defining probability paths directly on non-Euclidean manifolds (scale, rotation, opacity), ensuring all intermediate states are valid.
Abstract
We present Depth Anything V4 (DAV4), a framework for dynamic 4D scene reconstruction from monocular video. Our key contribution is the application of Riemannian Flow Matching (RFM) to 4D Gaussian Splatting parameters, defining probability paths directly on non-Euclidean manifolds (scale, rotation, opacity), ensuring all intermediate states are valid. Through controlled experiments, we isolate RFM's contribution from test-time optimization (TTO) and pre-training. A deterministic MLP baseline with the same data, architecture, and TTO achieves F-score 0.762; RFM achieves 0.806 - the +0.044 gain is RFM's isolated contribution. We provide corrected computational cost analysis: pre-training is 360 GPU-hours, amortizing for large-scale deployment (over 10,000 scenes). Uncertainty is quantified via Negative Gaussian Log-Likelihood and Expected Calibration Error. DAV4 outperforms prior Depth Anything models and per-scene 4D-GS on dynamic reconstruction and novel-view synthesis, while using no human-annotated depth labels as training losses.
Map-Det3D is an online multi-view 3D object detection model that brings detection directly into a 3D space reconstructed from RGB, suggesting that training reconstruction priors for detection is a practical route to stable metric 3D detection from monocular video.
Yung-Hsu Yang, Luigi Piccinelli, S. R. Bulò et al.· 0 citations
Render--match--PnP relocalization establishes correspondences between query image pixels and 3D map points for camera pose recovery, but their potential to support dense depth estimation is often overlooked. To exploit this geometric information, we present RIDE, which estimates dense metric depth from a robot's RGB stream. Given a metrically scaled 3D Gaussian Splatting (3DGS) model, RIDE combines sparse metric depth observations derived from PnP-RANSAC inlier correspondences with the geometric prior of a pretrained video-depth model. To handle uneven and intermittent observations, it integrates global and local depth correction with temporal memory, supporting depth estimation through short observation gaps after metric scale initialization. Trained on public RGB-D videos, RIDE is evaluated on robot sequences without fine tuning. Experiments show improved depth accuracy and temporal consistency over scale-only calibration, demonstrating how localization geometry can support both pose recovery and dense robot perception.
Jia-Rong Lian, Zhen-Hua Xiao, Zhao-Yang Zhang et al.· 0 citations
3D Gaussian Splatting has become a de facto scene representation for novel view synthesis, yet robustly learning 3D Gaussian primitives from visual input remains challenging. Standard optimization relies on gradient-based updates, but a common issue is the gradient vanishing phenomenon: a pixel far from a Gaussian primitive often has diminishing gradient magnitudes to influence primitive attributes, resulting in suboptimal scene reconstruction. In this paper, we propose a method to address gradient vanishing with a piecewise truncated gradient formulation that improves the optimization stability and robustness to initializations. We show that our method consistently improves 3D Gaussian Splatting with random and COLMAP initializations while being generalizable across static and dynamic Gaussian Splatting. As a by-product, we also examine the limitations of current benchmarks for dynamic scenes, and introduce a novel dataset for benchmarking dynamic Gaussian Splatting using synthetic 3D scenes. We demonstrate the effectiveness of our method in both static and dynamic settings for the public benchmarks and our proposed dataset.
Théo Morales, Nhat-Quynh Le-Pham, Robin Atkins et al.· 0 citations
Extremely sparse-view 3D Gaussian Splatting (3DGS) often suffers from depth drift, incorrect occlusions, and floating artifacts because supervision is unavailable between training cameras. We propose a stability-guided relative geometry distillation (SG-RGD) framework that extends geometric supervision to interpolated pseudo-views while reducing the influence of unreliable regions. Perturbation-based Pseudo-view Stability (PVS) estimates pixel-wise continuous stability weights from appearance differences between normal and mildly scale-perturbed renderings at the same pseudo-camera. Relative geometry distillation (RGD) uses these weights to modulate scale- and shift-invariant correlation alignment between Gaussian-rendered depth and relative depth predicted by a frozen monocular teacher. Progressive Multi-scale Pseudo-depth Curriculum (PMPC) introduces this supervision in a coarse-to-fine manner. On the nine-scene Mip-NeRF 360 three-view benchmark, SG-RGD improves PSNR from 10.6955 to 10.9600 dB and SSIM from 0.1422 to 0.2150 while reducing LPIPS from 0.7146 to 0.7100 and AVGE from 0.3849 to 0.3716 compared with DNGaussian. Experiments on Mip-NeRF 360 six-view and NeRF Synthetic eight-view further characterize performance across view coverage and data distributions, while ground-truth geometry diagnostics show that lower stability weights tend to correspond to larger geometric errors. All proposed modules are training only, preserving the standard 3DGS inference pipeline.
Shuai Gao, Fan Zhou, Shi-Wei Shao· Electronics· 0 citations
UniQuery4R is presented, a query-conditioned framework that encodes a multi-frame clip once and selects the source view, target view, and continuous source-image coordinate only at decoding time via source-to-target cross-attention, and introduces a direction-magnitude parameterization of scene flow with separate supervision for moving and static points.
Tiancheng Chen, Sheng Tang, Wenhua Jin et al.· 0 citations
3D Gaussian Splatting (3DGS) enables real-time rendering of photorealistic scene representations from multiview images and has potential as a visualization layer for digital twins. However, photorealistic appearance and image-level metrics alone do not establish practical suitability. We evaluate standard 3DGS on six real-world datasets: wall, forest, piled-pier underside, steel-girder bridge, asphalt pavement, and outdoor sculpture. The framework combines held-out test-view metrics (PSNR, SSIM, and LPIPS), spatial error maps, region-of-interest (ROI) inspection, computational records, scale–opacity grouping, and Gaussian-to-MVS distance analysis. In the wall case study, extending adaptive density control from 7000 to 30,000 iterations nearly doubled Gaussian count and model size and increased training time by about 50%, with only modest test-view improvements. Scale–opacity classes did not reliably distinguish high- from low-error regions. Gaussian-to-MVS distance was more strongly associated with scale than learned opacity, and median local errors generally increased with distance, although distributions overlapped substantially. Across datasets, major structures were generally reproduced well, whereas sparse foliage, low-contrast repetitive textures, fine surface details, and distant background objects required ROI-level verification. The best test-view performance reached 37.11 dB PSNR, 0.952 SSIM, and 0.193 LPIPS. Practical evaluation should consider global quality together with local, computational, and geometric characteristics.
Tomohiro Mizoguchi· Italian National Conference...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.