Skip to content
Conference

Smartphone-Based Obstacle Distance Estimation for Assistive Navigation Using YOLO and Monocular Depth Fusion: Architecture and Preliminary Baseline Evaluation

Aug 2026 · International Conferences on Information Science and System · pp. 1-9 · 0 citations · 23 references

Abstract

Navigating complex environments presents significant challenges for the estimated 285 million visually impaired individuals worldwide, directly impacting their mobility and quality of life [1]. While traditional assistive systems rely on expensive and bulky hardware-based depth sensors, such as LiDAR or structured light cameras [2], recent advances in deep learning enable software-based depth reconstruction from a single monocular camera, eliminating the need for dedicated hardware [3]. This paper presents the architecture and preliminary baseline evaluation of a smartphone-based assistive navigation system that fuses YOLO object detection with MiDaS monocular depth estimation (MDE) on a commodity Android device. The system is designed around a tiered estimation pipeline: a fully operational MDE-fusion mode (Tier-1) that combines MiDaS relative depth with one-point metric calibration, and a YOLO-affine fallback mode (Tier-3) that provides geometry-based distance estimation without calibration. This paper reports the baseline evaluation of the YOLO-affine mode, which establishes the system’s minimum-capability performance floor and demonstrates architectural readiness for full MDE-fusion evaluation. The system employs YOLOv8 Nano and MiDaS Small—both running on the smartphone CPU via ONNX Runtime—and delivers auditory proximity alerts through a three-zone classification scheme (Near, Medium, Far). Preliminary evaluation on 71 observations across seven distances (0.5–4.0 m) on a mid- range Snapdragon 680 device achieves a mean absolute error of 0.130 m and a zone classification accuracy of 98.6%—a 58.3% MAE reduction over a monocular geometric baseline and a 6.7% mean relative error over the pedestrian-critical 1.0–3.0 m range. MiDaS depth inference was confirmed operational on all 71 frames, with relative depth values exhibiting monotonic distance ordering, establishing the prerequisite condition for Tier-1 metric calibration.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.