3D Gaussian Splatting for Large-Scale Remote Sensing: A PRISMA-Informed Scoping Review of Scalability, Geometric Reliability, and Benchmarking Across UAV/Aerial and Satellite Imagery
A platform-aware benchmark framework that jointly records visual fidelity, computational cost, metric geometry, product utility, failure behavior, and reproducibility metadata for UAV/aerial, satellite, and hybrid settings is proposed.
Abstract
3D Gaussian Splatting (3DGS) offers efficient explicit rendering, but large-scale remote-sensing use remains fragmented across UAV/aerial photogrammetry, satellite reconstruction, large-scene scaling, surface modeling, and geospatial evaluation. We present a Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA)-informed scoping review based on 55 core studies identified through Web of Science, Scopus, IEEE Xplore, and supplementary searches completed on 3 June 2026. A faceted taxonomy organizes the literature by platform, sensor model, scalability strategy, and geometric supervision. The synthesis shows that partitioning, hierarchy, compression, and feed-forward inference improve scalability but do not guarantee metric geometry. Reliable deployment additionally requires sensor-consistent projection, geometric or georeferencing constraints, explicit supervision labels, and product-level evaluation. In control-point-free settings, internal consistency should be distinguished from independently validated accuracy. We therefore propose a platform-aware benchmark framework that jointly records visual fidelity, computational cost, metric geometry, product utility, failure behavior, and reproducibility metadata for UAV/aerial, satellite, and hybrid settings.
Abstract. Satellite imagery offers a distinct advantage in Earth observation by providing expansive coverage and enabling the monitoring of inaccessible regions without physical on-site intervention, serving as a significantly more cost-effective and scalable alternative to traditional aerial or ground-based surveys. The task of 3D reconstruction from multi-view satellite images has therefore been a pivotal point of research at the intersection of photogrammetry and remote sensing. Recently, novel-view synthesis techniques such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have accelerated the accuracy and speed of topographic modeling. Among these, Earth Observation Gaussian Splatting (EOGS) has emerged as a state-of-the-art approach by adapting 3DGS to handle the unique geometric and radiometric characteristics of satellite data, including Rational Polynomial Coefficients (RPCs) and varying solar conditions. However, the standard EOGS pipeline relies on stochastic initialization, where Gaussians are distributed uniformly within a volumetric bounding box, leading to high computational overhead and dependency on aggressive pruning that can inadvertently remove critical geometric features, particularly in areas with complex urban structures. To address these limitations, we propose Bundle-Adjusted Initialization for Earth Observation Gaussian Splatting, which leverages sparse point clouds from bundle adjustment as geometric priors for Gaussian initialization. Combined with an adaptive densification strategy, our method achieves faster convergence and improved DSM accuracy on the DFC2019 dataset compared to the EOGS baseline.
Jiyong Kim, Shuang Song, Rongjun Qin· The International Archives o...· 0 citations
This article presents a scalable and stable 3-D Gaussian splatting (3DGS)-simultaneous localization and mapping (SLAM) framework for efficient large-scale orthophoto generation. Unlike conventional SfM- or SLAM-based pipelines that rely on geometry-driven stitching, we reformulate orthophoto generation as a rendering-based mapping problem under a unified 3DGS-SLAM paradigm. However, directly applying 3DGS-SLAM to aerial mapping suffers from critical challenges, including uncontrolled memory growth, optimization instability, and slow convergence in large-scale UAV scenarios. To address these issues, we introduce a unified framework that jointly enforces memory-constrained representation, stabilized optimization, and accelerated convergence during incremental mapping. Specifically, we design an online gradient-driven mechanism to regulate Gaussian evolution, a viscous velocity regularization to stabilize optimization dynamics, and a geometry-aware homography-guided densification strategy to accelerate convergence under planar scene priors. Furthermore, by aligning GNSS with the SLAM system, our framework enables globally consistent orthographic rendering, producing geometrically and geographically consistent orthophotos in a unified coordinate frame. Extensive experiments on multisource UAV datasets, including those equipped with ground control points (GCPs), demonstrate that the proposed method achieves high absolute metric accuracy, along with superior efficiency, stability, and visual fidelity, enabling practical real-time incremental orthophoto generation for large-scale aerial environments.
Xiao Zhang, Hongbin Dong, Xiaozhou Zhu et al.· IEEE Transactions on Geoscie...· 0 citations
3D Gaussian Splatting (3DGS) has gained prominence in autonomous driving and robotics for its rendering efficiency and high-fidelity reconstruction capabilities. However, incremental 3DGS map construction at large scales remains challenging due to sensor sparsity and computational constraints. In this paper, we propose LV-GS SLAM, a novel system that integrates LiDAR and visual data for incremental, large-scale reconstruction with real-time tracking. This system provides fast and robust pose estimation while also enabling photorealistic rendering. We first employ a LiDAR odometry frontend that processes 30 Hz LiDAR inputs to provide robust initial poses. In our implementation, the complete LiDAR tracking pipeline runs at 15–19 Hz, while the mapping module performs incremental optimization on selected keyframes. To address the sparsity-induced surface discontinuity in conventional LiDAR-based reconstruction, we propose a novel depth propagation approach that initializes 3D Gaussian primitives using dense depth maps, achieving faster PSNR convergence with 9× fewer optimization iterations compared to direct LiDAR initialization. Furthermore, we develop a keyframe-based submap management framework that dynamically adjusts memory allocation based on both primitive density and inter-frame overlap ratio, effectively preventing GPU memory overflow. Our system has been validated on the KITTI dataset, achieving superior rendering quality compared with representative reproducible baselines. We further validate the robustness of the system on a quadruped robot platform, demonstrating satisfactory performance in both pose estimation and high-fidelity reconstruction.
Abstract. Monocular depth estimation (MDE) has reached notable maturity in computer vision, yet its application to UAV-based architectural heritage documentation remains underexplored. This study assesses whether the depth foundation model Depth Anything V2 can be transferred from terrestrial to aerial imagery. The analysis relies on MDE4BH, a benchmark of over 3,000 UAV images covering ten heterogeneous heritage scenarios (urban areas, façades, towers, villas, domes, and archaeological sites). Masked photogrammetric depth maps serve as metric reference for calibration, validation, and supervised retraining. Two baseline configurations are evaluated: a relative model with scene-specific linear rescaling and the direct application of the metric model. The rescaled relative model shows acceptable performance in several subsets, whereas the metric model exhibits systematic bias, weak consistency, and scale collapse due to domain shift between terrestrial training data and aerial acquisition geometry. To address these limitations, a two-step fine-tuning strategy is introduced, focusing on the decoder and regression head. The first stage uses mainly oblique UAV images; the second integrates oblique and nadir views to improve viewpoint generalization. The adapted model significantly reduces bias and enhances metric stability across the benchmark. However, residual errors remain spatially structured, with clustering and recurrent artefacts near object boundaries, multi-level roofs, and radiometrically heterogeneous surfaces. Although accuracy is still insufficient for demanding metric applications, the results support the use of MDE as a complementary source for thematic interpretation, scene understanding, robotics, navigation, and related tasks where strict geometric precision is not required.
F. Chiabrando, Francesca Gallitto, A. Lingua et al.· The International Archives o...· 0 citations
True digital orthophoto maps (TDOMs) serve as foundational geospatial products for applications in land surveying, urban planning, and emergency management. Conventional TDOM generation relies on differential correction, often resulting in cartographic artifacts such as geometric discontinuities, radiometric inconsistencies, and linear feature misalignments. Although recent methods leverage 3D Gaussian splatting to bypass differential correction, their computational demands hinder real-world deployment. To address these limitations, we propose FastPro-Gaussian, a novel framework enabling rapid high-quality TDOM generation on consumer-grade GPUs. Our contributions are threefold: block-based processing ensuring scalability for large-scale scenes; progressive densification stabilizing model optimization and reducing training iterations; and spherical-to-ellipsoidal Gaussian transformation, accelerating early-stage optimization of structural features (e.g., building edges). Experiments demonstrate that FastPro-Gaussian surpasses commercial solutions (ContextCapture, Metashape, and Pix4Dmapper) in rendering quality for building facades, edges, and roads. Compared to state-of-the-art methods, it achieves comparable TDOM fidelity with over two-fold acceleration in training time (notably more than two times faster than Tortho-GS). These gains in efficiency and efficacy confirm its strong potential for practical deployment in geospatial production pipelines.
Chao Yang, Yapeng Li, Feiyang Liu et al.· Photogrammetric Engineering...· 0 citations
Three-dimensional (3D) imaging systems, including depth cameras, LiDAR sensors, and multi-view scanning pipelines, often produce point clouds with noisy normals, outliers, sparse sampling, and non-uniform density, which can degrade downstream mesh reconstruction. Poisson surface reconstruction is lightweight and training-free, but its global implicit formulation is sensitive to unreliably oriented samples and fixed density-trimming thresholds. This paper presents RG-PSR, a reliability-guided enhancement framework for Poisson-family surface reconstruction from degraded 3D-imaging point clouds. RG-PSR estimates a deterministic per-point reliability score from local density regularity, spacing variation, and normal consistency, and propagates this score through conservative point filtering, reliability-guided normal refinement, adaptive density-reliability trimming, and structure-aware postprocessing. The main pipeline requires no manual labels, neural network training, or ground-truth meshes at inference time. Experiments on three groups of object meshes under five deterministic degradation types show that RG-PSR improves Poisson-family reconstruction under degraded inputs. Compared with fixed density-trimmed Poisson reconstruction, RG-PSR reduces the overall Chamfer-L1 from 0.0218 to 0.0172, improves F0.01 from 0.6618 to 0.6836, and reduces Artifact0.02 from 0.3090 to 0.2632. In the broader classical comparison, local triangulation methods achieve stronger point-wise accuracy, while RG-PSR yields the fewest connected components and the highest largest-component ratio. These results position RG-PSR as a practical reliability layer for coherent Poisson-family reconstruction rather than a universal replacement for all surface-reconstruction methods.
Na Liu, Fan Zhang, Jia-Wei Wang et al.· Journal of Imaging· 0 citations