Abstract. Accurate building footprint extraction is critical for applications ranging from population estimation to disaster management. Although optical imagery provides detailed spectral information, it often struggles with shadows, occlusions, and background clutter in dense urban environments. Lidar data, by contrast, offer precise elevation and structural attributes but face challenges such as variable point density and noise. This study integrates multispectral imagery from the U.S. Department of Agriculture (USDA) National Agriculture Imagery Program (NAIP) with lidar-derived feature height and intensity from the U.S. Geological Survey (USGS) 3D Elevation Program (3DEP) to improve footprint extraction using a U-Net–based deep learning model. A six-band input stack (RGB, near-infrared, height, intensity) was developed, normalized, and tiled for training and evaluation against Microsoft Global Building Footprints (GBF). Results from the Houston, TX test site show that the six-band model achieved a precision of 0.86, recall of 0.88, F1 score of 0.87, and Intersection-over-Union (IoU) of 0.76, consistently outperforming four-band baselines by reducing false positives while maintaining sensitivity. Predictions on withheld Houston tiles confirmed strong within-region generalization, yielded a precision of 0.78, recall of 0.81, F1 score of 0.79, and IoU of 0.66. Qualitative analysis further revealed limitations stemming from both training label quality and vegetation–building confusion. These findings demonstrate the complementary value of integrating spectral and structural information for robust building footprint extraction and how domain adaptation strategies can be used to enhance cross-regional transferability.
Jung-Kuan Liu, Rongjun Qin, S. Arundel et al.· ISPRS Annals of the Photogra...· 0 citations
Image-based 3D reconstruction is vital in many applications, such as digital twins, smart cities, machine vision, and autonomous driving. In recent years, it has undergone a paradigm shift, propelled by advancements in both conventional photogrammetry and deep learning. This review provides a comprehensive photogrammetric perspective on both conventional and learning-based techniques, a viewpoint that prioritizes geometric fidelity, robustness, handling of uncertainty, and suitability for real-world applications. We first systematically revisit the fundamentals of traditional pipelines: Structure from Motion (SfM), Multi-View Stereo (MVS), and surface reconstruction. The review then details recent progress in conventional methods, highlighting innovations in scalable and efficient SfM, specialized camera models for MVS, and robust surface reconstruction algorithms. Subsequently, we explore the transformative evolution brought by learning-based techniques, including deep SfM, learning-based MVS, differentiable rendering-based scene representation methods (NeRF, 3DGS), groundbreaking feed-forward 3D reconstruction models (e.g., DUSt3R, VGGT), and surface reconstruction including explicit and implicit methods. Emphasis is placed on evaluating whether learning-based approaches genuinely meet photogrammetric requirements such as metric accuracy and reliability, rather than optimizing solely for visual realism.
Furthermore, we conclude by identifying key challenges and research frontiers including generalization across domains, scalability to high-resolution imagery, real-time performance, and uncertainty quantification. By bridging the gap between classical photogrammetry and data-driven 3D vision, this work aims to guide future research toward robust, accurate, and certifiable 3D reconstruction systems suitable for engineering, industrial, and geospatial applications.
Xin Wang, Tengfei Wang, M. Hillemann et al.· PFG – Journal of Photogramme...· 0 citations
Abstract. Among the image-based methods, photogrammetry is a consolidated 3D reconstruction technique able to provide highly accurate metric products, widely exploited in many domains. Photogrammetry is, however, conditioned by the characteristics of the captured scene, with good performance in well-textured areas and limits when non-collaborative surfaces, such as reflective or transparent, are present. In such cases, the photogrammetric reconstruction is often affected by noise, incomplete geometry and artifacts, reducing its final reconstruction quality. In recent years, different AI-based reconstruction methods have emerged as alternative (or complementary) 3D reconstruction and rendering solutions. In particular, 3D Gaussian Splatting (GS) has demonstrated impressive capabilities in rendering photorealistic scenes in challenging situations with high visual fidelity. However, its application in large-scale scenarios or when highly accurate 3D metric products are required is still limited, due to the high computational resources needed and the intrinsic optimization of GS methods for photometric rendering quality. To address these bottlenecks, this work proposes a hybrid reconstruction pipeline, leveraging the strengths and benefits of each technique. The method exploits the accurate geometry of photogrammetry in well-textured regions and the GS capabilities to improve completeness and visual aspect in areas featuring non-collaborative surfaces. A fusion strategy is proposed to combine the two results into a single 3D model, presenting examples from two aerial and one terrestrial dataset.
Fabio Remondino, E. M. Farella, Gianluca Bertolasi et al.· The International Archives o...· 0 citations
Abstract. Satellite imagery offers a distinct advantage in Earth observation by providing expansive coverage and enabling the monitoring of inaccessible regions without physical on-site intervention, serving as a significantly more cost-effective and scalable alternative to traditional aerial or ground-based surveys. The task of 3D reconstruction from multi-view satellite images has therefore been a pivotal point of research at the intersection of photogrammetry and remote sensing. Recently, novel-view synthesis techniques such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have accelerated the accuracy and speed of topographic modeling. Among these, Earth Observation Gaussian Splatting (EOGS) has emerged as a state-of-the-art approach by adapting 3DGS to handle the unique geometric and radiometric characteristics of satellite data, including Rational Polynomial Coefficients (RPCs) and varying solar conditions. However, the standard EOGS pipeline relies on stochastic initialization, where Gaussians are distributed uniformly within a volumetric bounding box, leading to high computational overhead and dependency on aggressive pruning that can inadvertently remove critical geometric features, particularly in areas with complex urban structures. To address these limitations, we propose Bundle-Adjusted Initialization for Earth Observation Gaussian Splatting, which leverages sparse point clouds from bundle adjustment as geometric priors for Gaussian initialization. Combined with an adaptive densification strategy, our method achieves faster convergence and improved DSM accuracy on the DFC2019 dataset compared to the EOGS baseline.
Jiyong Kim, Shuang Song, Rongjun Qin· The International Archives o...· 0 citations
High-resolution satellite imagery demands three-dimensional (3D) reconstruction methods that deliver both speed and geometric accuracy. Recent adaptations of 3D Gaussian splatting (3DGS) to satellite imagery demonstrate strong efficiency, but reconstruction quality often degrades under diverse illumination across multi-date, high-altitude acquisitions (with small intersection angles), limiting applicability to remote sensing and vision tasks. We present SatSplat, the first framework to adapt 2D Gaussian splatting (2DGS) to satellite photogrammetry, with online camera adjustment. We approximated satellite cameras with an affine model and learned a minimal delta parameterization for in-splat camera refinement from dense observations. The formulation was implemented with a 2DGS scene representation. To handle time-varying shadows and illumination changes, we integrated geometric shadow mapping and per-camera color correction during training. Across the evaluated DFC2019 and IARPA2016 benchmark sites, SatSplat achieved strong geometric accuracy while significantly outperforming prior 3DGS-based baselines. On our processed DFC2019 benchmark, SatSplat reduced mean absolute error by 11.93% and peak video memory by 31% relative to the previous state of the art. Our approach enabled large-scale digital surface modeling with practical computational efficiency. The project page is available at https://gdaosu.github.io/satsplat.
Shuang Song, Jiyong Kim, Rongjun Qin· Photogrammetric Engineering...· 1 citation