Preprint
Sep 2026
VLMs Can Describe, But Not Measure: Object-Centric Scene Understanding for Robotic Manipulation
A VLM-driven, modular perception framework for scene understanding using off-the-shelf approaches that preserves strong semantic performance while substantially improving localization and depth estimation over direct VLM inference.
Enrico Saccon, Tommaso Faraci, Iñigo De La Ossa Zarzuelo et al.
· 0 citations