CitySTAR: Structured and Topology-Aware Reasoning for Open-Vocabulary Urban 3D Grounding
3D grounding aims to localize target entities in complex scenes from natural language and plays a fundamental role in embodied perception and spatial reasoning. However, existing approaches mostly rely on feature similarity or direct matching, making it difficult to connect natural-language intent with the implicit sem...