2026· IEEE Transactions on Automation Science and Engineering· Vol 23, pp. 14592-14602· 0 citations· 52 references
Computer Science
Abstract
Transparent objects, such as beakers, flasks, and other laboratory glassware, are commonly used in both laboratory and industrial settings. When light passes through these materials, it is refracted, which degrades image quality and impairs depth estimation of object surfaces. This often leads to failures in robotic perception of transparent targets, reducing task efficiency and potentially creating safety hazards. To address this problem, we propose a vision–language-guided Transparent Object Perception (vToP) method that combines cognitive augmentation with confidence-based depth completion. First, using structural priors and cognitive feedback from vision–language models, transparent objects are redesigned as deformable models, allowing them to adaptively deform and align with observed object contours. Next, regional confidence scores are designed to evaluate depth quality for each pixel and are then used as weighting factors to dynamically fuse depth estimation with cognitive 3D modeling, recovering missing depth information of transparent objects. Finally, we evaluate vToP on the public TransCG dataset and a newly constructed synthetic dataset, CSSG, designed for transparent object perception. Experimental results demonstrate that vToP significantly outperforms existing methods in transparent object depth recovery. Furthermore, vToP can be seamlessly integrated with grasp detection algorithms, and real-world experiments on transparent objects using a UR5 robotic arm validate its effectiveness. Note to Practitioners—Transparent objects are commonly encountered in laboratory and industrial settings. However, conventional RGB-D cameras often produce inaccurate depth estimates when capturing transparent or highly reflective objects. This limitation leads to failures in object recognition, grasping, and manipulation, as well as potential safety hazards. Therefore, a reliable solution for transparent object perception is critically important. The proposed method, vision–language-guided Transparent Object Perception (vToP), is inspired by biological visual cognition. It identifies and completes uncertain depth regions by integrating structural priors with cognitive feedback from vision–language models, thereby significantly improving the accuracy of transparent object perception. Experimental results demonstrate that vToP not only enhances depth recovery but also achieves high success rates in real-world grasping tasks, highlighting its strong potential for applications in intelligent laboratories and other automated environments.
OptiGeo is introduced, a bias-aware training framework that rehabilitates biased real supervision using a clean-geometry teacher and residual-trimmed alignment and outperforms substantially larger 300M-scale monocular models and billion-scale multi-view baselines on transparent-scene benchmarks, while remaining competi...
Mu-Xin Liu, Tian-Bo Liu, Jing Xia et al.· 1 citation
Robotic manipulation relies on perceiving the object region that affords an instructed interaction, yet often assumes that this region is visible and accessible. In cluttered scenes, however, a handle, blade, tip, or other task-relevant region may be occluded. This paper introduces a zero-shot affordance exposure frame...
Wei-Lin Yao, Peng-Wen Xiong, Ying Liu et al.· 2026 3rd International Confe...· 0 citations
Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in scene reconstruction, computational photography, and robotics, among others. Despite the field's maturity, recent models still struggle to generalize to out-of-distribution inputs and to produce sharp a...
Igor Pavlovic, Thiemo Wandel, Anton Obukhov et al.· 0 citations
Robotic operation in previously unseen environments requires both semantic understanding and reliable metric information. While vision--language models (VLMs) provide strong semantic capabilities, their geometric estimates remain less reliable. In this paper, we propose a VLM-driven, modular perception framework for sc...
Enrico Saccon, Tommaso Faraci, Iñigo De La Ossa Zarzuelo et al.· 0 citations
Visual odometry (VO) estimates camera motion from image sequences and is essential for robotics, autonomous driving, and AR/VR. Robust VO remains challenging because large viewpoint changes and strong parallax make reliable cross-frame motion cues difficult to capture, especially in the presence of visual disturbances...
Jun-Qi Bao, Qing-Ying Wu, Jun Huang et al.· IEEE Transactions on Instrum...· 0 citations
Observing objects grasped by a robot hand is challenging due to severe visual occlusions. Although in-hand manipulation can expose hidden surfaces, existing approaches often rely on predefined or open-loop reorientation strategies that do not explicitly target under-observed regions. We propose AURORA, an active 3D rec...