Monocular distance estimation: from geometric foundations and deep learning innovations to industrial deployment challenges
Abstract
Monocular distance estimation, a fundamental yet ill-posed problem in computer vision, has evolved from rigid geometric constraints to flexible deep learning paradigms. This review provides a comprehensive analysis of this transition, categorizing methodologies into traditional geometric-based, supervised, and self-supervised frameworks. We examine how traditional methods (e.g., size priors and ground-plane constraints) offer interpretability but struggle with environmental robustness. In the deep learning era, we detail the shift from CNN receptive field limitations to Transformer based global dependency modeling (e.g., DPT, MonoViT) and the mathematical progression from continuous regression to adaptive depth binning. A significant focus is placed on self-supervised mechanisms, specifically the "photometric consistency" assumption and landmark innovations like SfM-Learner’s joint optimization and Monodepth2’s auto masking strategy. Finally, we synthesize performance benchmarks on the KITTI and NYU Depth V2 datasets to highlight current bottlenecks—namely, scale ambiguity and domain shift. The review concludes that future industrial deployment on edge-computing platforms will rely on a synergy between lightweight network architectures and multi-sensor fusion.