Skip to content
Open access

Evaluating the Performance of 3D Vision Foundation Models for DSM Reconstruction from Satellite Images

Jul 2026 · ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences · 0 citations · 14 references

TL;DR

Overall, this study systematically quantifies the performance of 3D VFMs in satellite image-based 3D reconstruction, confirming their strong potential for high-resolution satellite applications and providing valuable insights for enhancing model robustness and generalization across complex urban and low-resolution environments.

Abstract

Abstract. Three-dimensional (3D) reconstruction from satellite imagery is a critical research topic in the fields of remote sensing and geoinformation science. Although 3D Vision Foundation Models (3D VFMs) have demonstrated remarkable performance in reconstructing natural scenes, their capability to handle high-resolution satellite imagery has not been systematically evaluated. This study presents a comprehensive assessment of seven representative 3D VFMs for satellite-based 3D reconstruction and integrates four point-cloud alignment strategies. Rigorous comparisons were conducted against high-precision LiDAR-derived Digital Surface Models (DSMs) using two publicly available multi-view satellite datasets–WHU-TLC and MVS3D. The results show that Depth Anything V2 (DAV2) combined with an affine alignment strategy achieves the best overall performance among the evaluated methods. On the MVS3DM dataset, the reconstructed DSM achieves a Median Absolute Error(MedAE) of 1.693 m, a Root Mean Square Error (RMSE) of 3.649 m, and competitive reconstruction accuracy compared with several traditional photogrammetric pipelines. In contrast, on the lower-resolution WHU-TLC dataset, all 3D VFMs exhibited notable performance degradation, and the reconstructed results showed limited practical value, revealing persistent generalization challenges for current models in low-resolution scenarios. Overall, this study systematically quantifies the performance of 3D VFMs in satellite image-based 3D reconstruction, confirming their strong potential for high-resolution satellite applications and providing valuable insights for enhancing model robustness and generalization across complex urban and low-resolution environments.

Read PDF

Similar papers

Open access Jul 2026

Using NeRFs for UAV-based 3D reconstruction of complex scenes: A comparison to MVS

Abstract. High-resolution 3D documentation of cultural heritage sites is essential for their preservation. While terrestrial laser scanning (TLS) remains the gold standard, it is often cost-intensive compared to photogrammetry. This study evaluates three image-based reconstruction techniques, Multi-View Stereo (MVS), Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS), by applying them to a complex scene featuring a chapel and its surrounding vegetation, sensed from an uncrewed aerial vehicle (UAV). A hybrid TLS/MVS model provides a high-accuracy reference. Using identical interior and exterior camera parameters of the 105 UAV-acquired images, we generate dense point clouds with all methods and assess geometric accuracy and completeness using the M3C2 algorithm. Results show that MVS achieves superior accuracy (standard deviation of all M3C2 distances: MVS = 0.11 m, NeRF = 0.15 m), whereas NeRF attains up to 20% higher completeness, particularly in low-texture and vegetation-occluded regions. The 3DGS point cloud was deemed too sparse and was therefore not used for further analysis. The study highlights the potential of NeRFs to recover partially occluded or sparsely textured geometries that are challenging for MVS and suggests a complementary use of both approaches for cost-efficient documentation of cultural heritage.

Frederik Schulte, P. Akwensi, L. Winiwarter · 0 citations
Review Open access Jul 2026

Bundle-Adjusted Initialization for Efficient Earth Observation Gaussian Splatting

Abstract. Satellite imagery offers a distinct advantage in Earth observation by providing expansive coverage and enabling the monitoring of inaccessible regions without physical on-site intervention, serving as a significantly more cost-effective and scalable alternative to traditional aerial or ground-based surveys. The task of 3D reconstruction from multi-view satellite images has therefore been a pivotal point of research at the intersection of photogrammetry and remote sensing. Recently, novel-view synthesis techniques such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have accelerated the accuracy and speed of topographic modeling. Among these, Earth Observation Gaussian Splatting (EOGS) has emerged as a state-of-the-art approach by adapting 3DGS to handle the unique geometric and radiometric characteristics of satellite data, including Rational Polynomial Coefficients (RPCs) and varying solar conditions. However, the standard EOGS pipeline relies on stochastic initialization, where Gaussians are distributed uniformly within a volumetric bounding box, leading to high computational overhead and dependency on aggressive pruning that can inadvertently remove critical geometric features, particularly in areas with complex urban structures. To address these limitations, we propose Bundle-Adjusted Initialization for Earth Observation Gaussian Splatting, which leverages sparse point clouds from bundle adjustment as geometric priors for Gaussian initialization. Combined with an adaptive densification strategy, our method achieves faster convergence and improved DSM accuracy on the DFC2019 dataset compared to the EOGS baseline.

Jiyong Kim, Shuang Song, Rongjun Qin · 0 citations
Review Open access Jul 2026

Research on digital modeling of ancient building appearances based on UAV three-dimensional reconstruction

A digital exterior-modeling method based on unmanned aerial vehicle (UAV) oblique photogrammetry and point cloud processing is proposed. The difficulty encountered by conventional surveying methods in comprehensively acquiring spatial information from elevated and occluded areas of historic buildings is addressed by this method. The Jiuyun Fangding monument was selected as the study object. Images were acquired using a multi-altitude, multi-angle, layered circumferential flight strategy. A three-dimensional (3D) model was then reconstructed through feature matching, camera-pose estimation, multi-view stereo matching, point cloud registration, and texture mapping. Four representative dimensions were selected, and the measurements obtained from the reconstructed model were compared with field measurements. The results indicated that: (1) the absolute errors of the four dimensions ranged from 0.006 to 0.039 m, with the maximum spacing between the load-bearing columns exhibiting the largest absolute error of 0.039 m; (2) the relative errors ranged from 0.63% to 1.81%, with the width of the load-bearing column exhibiting the largest relative error of 1.81%; and (3) the reconstructed model provided a relatively complete representation of the overall architectural form and the principal structural components. The proposed method can therefore provide technical support for the digital documentation, 3D visualization, and dimensional verification of historic buildings.

Yuecheng Yin · 0 citations
Open access Jul 2026

ESGS: A 3D Reconstruction Method for the Martian Surface Based on Optical Remote Sensing Images

Mars exploration is an advanced field of global deep space exploration. Accurate three-dimensional reconstruction of the Martian surface topography is very important for autonomous navigation, scientific target recognition, and operation planning. In order to meet the analysis requirements of the Martian surface scene, this paper proposes an explicit surface-geometry-constrained Gaussian splatting (ESGS) method. Firstly, this method includes a normal and depth prior estimation network (NDN) that generates normal and depth priors from Martian surface image data, thereby promoting the fusion of semantic and multi-view contextual information to enhance the geometric accuracy of 3D reconstruction of the Martian surface. Secondly, we designed the Gaussian parameter-based deformable fusion network (GPDFN) to fuse multi-receptive-field feature information. Finally, we collected Martian surface remote sensing images from NASA, constructed a Martian surface 3D reconstruction dataset named Mars_3D using the COLMAP method, annotated depth and normal labels for its seven real-world scenes and two Blender-generated scenes, and conducted comparative experiments with eight excellent algorithms on this dataset to validate the effectiveness of our method in 3D reconstruction of the Martian surface using remote sensing images. Experiments show that the average SSIM of the ESGS method in this article is 0.6946, PSNR is 23.40 dB, and LPIPS is 0.253 on the Mars_3D dataset, demonstrating superior overall performance compared to all other models and enhancing the quality of 3D reconstruction of the Martian surface.

Qinghe Guan, Y. Liu, Lei Chen et al. · 0 citations
Open access Aug 2026

Monocular Depth Estimation from UAV Images for 3D Documentation of Architectural Heritage: A Depth Anything V2-Based Approach

Abstract. Monocular depth estimation (MDE) has reached notable maturity in computer vision, yet its application to UAV-based architectural heritage documentation remains underexplored. This study assesses whether the depth foundation model Depth Anything V2 can be transferred from terrestrial to aerial imagery. The analysis relies on MDE4BH, a benchmark of over 3,000 UAV images covering ten heterogeneous heritage scenarios (urban areas, façades, towers, villas, domes, and archaeological sites). Masked photogrammetric depth maps serve as metric reference for calibration, validation, and supervised retraining. Two baseline configurations are evaluated: a relative model with scene-specific linear rescaling and the direct application of the metric model. The rescaled relative model shows acceptable performance in several subsets, whereas the metric model exhibits systematic bias, weak consistency, and scale collapse due to domain shift between terrestrial training data and aerial acquisition geometry. To address these limitations, a two-step fine-tuning strategy is introduced, focusing on the decoder and regression head. The first stage uses mainly oblique UAV images; the second integrates oblique and nadir views to improve viewpoint generalization. The adapted model significantly reduces bias and enhances metric stability across the benchmark. However, residual errors remain spatially structured, with clustering and recurrent artefacts near object boundaries, multi-level roofs, and radiometrically heterogeneous surfaces. Although accuracy is still insufficient for demanding metric applications, the results support the use of MDE as a complementary source for thematic interpretation, scene understanding, robotics, navigation, and related tasks where strict geometric precision is not required.

F. Chiabrando, Francesca Gallitto, A. Lingua et al. · 0 citations
Open access Aug 2026

Construction Method of Multimodal 4D Imaging Radar Dataset for Three-Dimensional Traffic Scenes

The latest generation of 4D imaging radar demonstrates significant potential in autonomous driving environmental perception, leveraging its capability to provide target elevation data and dense point clouds. This paper introduces a complete method for constructing a multimodal 4D imaging radar dataset for three-dimensional traffic scenes. It illustrates the hardware and software configurations of the data-acquisition vehicle. Methods including multi-sensor coordination, parameter calibration, timestamp synchronization and spatial datum synchronization are proposed. And eight typical three-dimensional traffic scenarios are designed, such as rainy weather environments, dense heterogeneous targets, enclosed tunnels, high-speed cut-in of multiple vehicles, multi-layered stereoscopic structures and edge working condition reproduction. In addition, this paper puts forward a frame-by-frame processing method for high-resolution images and point cloud data collected by the high-definition camera-LiDAR-4D imaging radar collaborative system. A large model-based 3D annotation method for multiple types of targets is proposed, generating a spatio-temporal sequence-optimized four-dimensional annotation sequence, and finally constructs a complete and high-quality multimodal 4D imaging radar dataset for three-dimensional traffic scenes. The results show that the constructed dataset enables the synchronization of timestamps and spatial coordinate systems. The large model can achieve high-precision 3D annotation for the four predefined target types. The dataset contains 11,400 frames of data from high-definition cameras, LiDAR, and 4D imaging radar, with 131,642 labels. This study will provide reliable fundamental support for the training and verification of 4D imaging radar perception algorithms, vehicle decision-making and planning in complex scenarios, and multi-sensor fusion technologies.

Zhuanzhuan Zhao, Xin Zhang, Shengyu Yan et al. · 0 citations