Skip to content
Preprint

OptiGeo: Efficient Monocular Geometry for Embodied Perception in Optically Challenging Scenes

Aug 2026 · 1 citation · 64 references
Computer Science

TL;DR

OptiGeo is introduced, a bias-aware training framework that rehabilitates biased real supervision using a clean-geometry teacher and residual-trimmed alignment and outperforms substantially larger 300M-scale monocular models and billion-scale multi-view baselines on transparent-scene benchmarks, while remaining competitive on general zero-shot depth and boundary sharpness.

Abstract

Monocular depth estimation has achieved strong open-domain generalization, yet reliable robotic deployment remains difficult in transparent, reflective, and specular environments, where depth sensors often produce missing or biased depth. Existing methods often handle such optical failures with scene-specific preprocessing, auxiliary modules, or post-hoc fine-tuning. While effective in constrained settings, these designs increase architectural redundancy and can over-specialize general geometry models to narrow optical scenarios. We revisit this problem as a localized failure mode within base-model training and identify sensor-induced supervision bias as a key bottleneck: models inherit sensor failure patterns from biased real-depth supervision in optically challenging regions. We then introduce OptiGeo, a bias-aware training framework that rehabilitates biased real supervision using a clean-geometry teacher and residual-trimmed alignment. We redefine transparency-targeted rendering as a compact source of clean optical geometry, rather than a large domain-specific fine-tuning set. With only a small targeted rendering set, OptiGeo learns the geometric structure of transparent objects and regions, correcting local geometry distortions that real sensors cannot reliably supervise. Despite only 30M parameters, OptiGeo outperforms substantially larger 300M-scale monocular models and billion-scale multi-view baselines on transparent-scene benchmarks, while remaining competitive on general zero-shot depth and boundary sharpness. Real-world navigation cases further validate its practicality as an efficient perception module in optically challenging scenes.

View source

Similar papers

#machine learning Preprint Sep 2026

Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation

Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in scene reconstruction, computational photography, and robotics, among others. Despite the field's maturity, recent models still struggle to generalize to out-of-distribution inputs and to produce sharp a...

Igor Pavlovic, Thiemo Wandel, Anton Obukhov et al. · 0 citations
Preprint Aug 2026

LiteMVS: Efficient Multi-View Stereo with Foundation Distillation and Expert Aggregation

LiteMVS is a lightweight multi-view depth estimation model that integrates plane-sweep geometric reasoning with strong monocular semantic and structural priors and employs a Mixture-of-Experts (MoE) formulation to enable adaptive geometric aggregation across depth hypotheses.

Tian-Bao Zhang, Zeyu Liu, Shuyu Wu et al. · 0 citations
2026

vToP: VLM-Guided Transparent Object Perception for Robotic Manipulation

Transparent objects, such as beakers, flasks, and other laboratory glassware, are commonly used in both laboratory and industrial settings. When light passes through these materials, it is refracted, which degrades image quality and impairs depth estimation of object surfaces. This often leads to failures in robotic pe...

Kaixin Huang, J. Zhuang, Sichao Ye et al. · 0 citations
Conference Aug 2026

Monocular distance estimation: from geometric foundations and deep learning innovations to industrial deployment challenges

It is concluded that future industrial deployment on edge-computing platforms will rely on a synergy between lightweight network architectures and multi-sensor fusion and self-supervised frameworks.

Zi-Kang Fan, Zi-Hao Xiang, Jiang-Sheng Liu · 0 citations
Review Sep 2026

Monocular Depth Estimation from a Single Image: Progress and Opportunities

Monocular depth estimation has long stood as a fundamental challenge in computer vision, enabling a wide range of applications including 3D reconstruction, robotics, autonomous driving, and augmented reality. This survey traces the field's evolution from early learning-based methods to the emergence of transformative f...

Mu-Xin Liu, Xiaoyang Lyu, Yang-Tian Sun et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.