VLMs Can Describe, But Not Measure: Object-Centric Scene Understanding for Robotic Manipulation
Robotic operation in previously unseen environments requires both semantic understanding and reliable metric information. While vision--language models (VLMs) provide strong semantic capabilities, their geometric estimates remain less reliable. In this paper, we propose a VLM-driven, modular perception framework for sc...