Skip to content

Diffusion-Driven Multiview Depth Estimation With Epistemic and Morphological Priors for Urban Scene Reconstruction

2026 · IEEE Transactions on Geoscience and Remote Sensing · Vol 64, pp. 5633615-5633615 · 0 citations · 47 references

Abstract

Urban scene reconstruction plays a critical role in applications such as autonomous navigation and digital twin systems. As a key technique for scene reconstruction, multiview stereo (MVS) has made significant progress in general scenarios. However, it remains fragile in complex urban environments. Challenges such as weak textures, dense occlusion, depth adhesions, and illumination variations severely impair the effectiveness of traditional photometric and geometric consistency constraints, leading to unstable depth estimation and topological artifacts. To address these issues, we propose StableMVS, a robust depth estimation method that integrates semantic priors, structural cues, and perturbation-aware learning. Specifically, epistemic-structural prior infusion introduces cognitive priors and semantic layout from large vision models to enhance depth inference in ambiguous regions. Morphology-encoded depth field optimization (MeDF) leverages boundary-aware morphology to suppress distortions and topological artifacts near depth discontinuities. Furthermore, diffusion-propelled feature fidelity learning adopts a disturbance–restoration paradigm that improves feature stability under real-world perturbations such as lighting shifts and sensor noise. These components collectively form a unified framework that bridges global understanding with local consistency, yielding structurally coherent and resilient depth estimation across challenging urban scenarios. Experiments on multiple public datasets show that StableMVS achieves competitive accuracy under low-texture, occlusion, and illumination variation conditions, validating its effectiveness for real-world MVS applications. The code is available at https://github.com/IMOP-lab/StableMVS

View source

Similar papers

Preprint Sep 2026

D3GS: Depth, DINO, and RGB Diffusion Co-Guided 3D Gaussian Splatting for Sparse-View Reconstruction

Novel view synthesis from sparse inputs remains challenging for 3D Gaussian Splatting (3DGS) due to ambiguous geometry, cross-view inconsistency, and missing details in under-constrained regions, resulting in degraded reconstruction and unstable rendering. To tackle these issues, we propose D$^{3}$GS, a Depth-DINO-Diff...

Yun-Qi Gao, Zhan-Feng Liao, Han-Zhang Tu et al. · 0 citations
Preprint Sep 2026

VDGS: Visibility-Driven Large-Scale 3D Gaussian Splatting for Aerial Scene Reconstruction

VDGS introduces visibility-driven statistics for scene anchors to quantify supervision strength and is leveraged for scene partitioning and for gradient compensation in under-optimized regions, thereby promoting balanced optimization across different regions.

Hao-Lin Yu, Jia-Dong Tang, Yi-Xian Wang et al. · 0 citations
#machine learning Preprint Sep 2026

Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation

Marigold V2 achieves state-of-the-art results when applied to other dense regression tasks, such as surface normals estimation and intrinsic image decomposition, and achieves state-of-the-art results when applied to other dense regression tasks, such as surface normals estimation and intrinsic image decomposition.

Igor Pavlovic, Thiemo Wandel, Anton Obukhov et al. · 1 citation · ⚡1
Open access 2026

OccluFree: Occlusion-Aware Restoration for Amodal Building Windows and Doors Segmentation

Urban applications, such as facade reconstruction, digital twins, building analysis, and semantic 3-D city modeling benefit from reliable and geometrically consistent representations of window and door (WD) elements. Many image-based facade reconstruction pipelines first detect WD elements in street-level imagery befor...

Abbas Salehitangrizi, S. Jabari, Yun Zhang · 0 citations
Preprint Oct 2026

DecomVoxel: Harnessing 3D-Native Priors with Guided In-situ Denoising Optimization for Decompositional Scene Reconstruction

Decompositional scene reconstruction aims to reconstruct high-quality objects and background, yet existing methods still struggle with the level of quality under heavy occlusions. While generative priors offer a potential solution, 2D image-based priors often suffer from multi-view inconsistency due to a lack of 3D awa...

Jun-Feng Ni, Zi-Rui Zhou, Yi-Xin Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.