Abstract. Monocular depth estimation (MDE) infers depth from a single image, offering significant advantages in computational efficiency and memory consumption compared to conventional Multi-View Stereo (MVS) methods. However, most MDE methods suffer from poor multi-view geometric consistency, which limits their application to photogrammetric 3D reconstruction. To address this issue, this paper employs sparse point clouds of Structure-from-Motion (SfM) as extra geometric constraints and proposes a framework that achieves photogrammetric 3D reconstruction using off-the-shelf learning-based MDE models without the need for additional fine-tuning. Specifically, when SfM priors are available during inference, globally geometrically consistent depth maps can be directly predicted. Otherwise, the estimated monocular depths are aligned to a consistent scale using SfM results via a post-correction step. The resulting depth maps are then fused using a truncated signed distance function (TSDF) to generate dense 3D reconstructions. Experiments on photogrammetric datasets demonstrate that the proposed framework effectively improves geometric consistency across depth maps and enables high-quality scene reconstruction. In addition, we systematically analyze the impact of key parameters in depth inference and fusion, including depth map resolution, voxel size, denoising steps, and ensemble size, on reconstruction performance, and further explore the potential of MDE for photogrammetric 3D reconstruction.
Chunyu Dou, Yifei Yu, Xin Wang et al.· The International Archives o...· 0 citations
Image-based 3D reconstruction is vital in many applications, such as digital twins, smart cities, machine vision, and autonomous driving. In recent years, it has undergone a paradigm shift, propelled by advancements in both conventional photogrammetry and deep learning. This review provides a comprehensive photogrammetric perspective on both conventional and learning-based techniques, a viewpoint that prioritizes geometric fidelity, robustness, handling of uncertainty, and suitability for real-world applications. We first systematically revisit the fundamentals of traditional pipelines: Structure from Motion (SfM), Multi-View Stereo (MVS), and surface reconstruction. The review then details recent progress in conventional methods, highlighting innovations in scalable and efficient SfM, specialized camera models for MVS, and robust surface reconstruction algorithms. Subsequently, we explore the transformative evolution brought by learning-based techniques, including deep SfM, learning-based MVS, differentiable rendering-based scene representation methods (NeRF, 3DGS), groundbreaking feed-forward 3D reconstruction models (e.g., DUSt3R, VGGT), and surface reconstruction including explicit and implicit methods. Emphasis is placed on evaluating whether learning-based approaches genuinely meet photogrammetric requirements such as metric accuracy and reliability, rather than optimizing solely for visual realism.
Furthermore, we conclude by identifying key challenges and research frontiers including generalization across domains, scalability to high-resolution imagery, real-time performance, and uncertainty quantification. By bridging the gap between classical photogrammetry and data-driven 3D vision, this work aims to guide future research toward robust, accurate, and certifiable 3D reconstruction systems suitable for engineering, industrial, and geospatial applications.
Xin Wang, Tengfei Wang, M. Hillemann et al.· PFG – Journal of Photogramme...· 0 citations