Aug 2026· Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2· pp. 10246-10257· 0 citations· 65 references
TL;DR
This work proposes UGE, a two-stage training strategy that progressively and stably aligns images, text, and spatial structures by combining instruction-guided contrastive learning with graph-based spatial encoding, and introduces \dataset, a spatially grounded dataset that anchors street-view images to structured spatial graphs and provides graph-aligned supervision via spatial reasoning paths and spatial context captions.
Abstract
Learning transferable multimodal embeddings for urban environments is challenging because urban understanding is inherently spatial, yet existing datasets and benchmarks lack explicit alignment between street-view images and urban structure. We introduce \dataset, a spatially grounded dataset that anchors street-view images to structured spatial graphs and provides graph-aligned supervision via spatial reasoning paths and spatial context captions, exposing distance, directionality, connectivity, and neighborhood context beyond image content. Building on UGData, we propose UGE, a two-stage training strategy that progressively and stably aligns images, text, and spatial structures by combining instruction-guided contrastive learning with graph-based spatial encoding. Lastly, we introduce UGBench, a comprehensive benchmark to evaluate how spatially grounded embeddings support diverse urban understanding tasks, including geolocation ranking, image retrieval, urban perception, and spatial grounding. In particular, UGE is built on multiple state-of-the-art VLM backbones: Qwen2-VL, Qwen2.5-VL, Phi-3-Vision, and LLaVA1.6-Mistral, with fixed-dimensional spatial embeddings trained under LoRA tuning. UGE built upon Qwen2.5-VL-7B backbone achieves up to 44% improvement in image retrieval and 30% in geolocation ranking on training cities, and over 30% and 22% gains respectively on held-out cities, demonstrating the effectiveness of explicit spatial grounding for spatially intensive urban tasks. The code and datasets are available at https://github.com/Jaygagaga/UGE/tree/main.
Geospatial representation learning from satellite imagery is a fundamental problem for large-scale urban analysis and real-world applications. Despite recent advances, current methods struggle with cross-region generalization and semantic interpretability due to their reliance on region-specific auxiliary data and the...
Yutian Jiang, Jiabo Liu, Xixuan Hao et al.· 1 citation
The results show that distributionally-regularized joint embedding architectures can be successful on challenging city-scale 3D scenes, and that transfer improves when self-supervision is designed for the capture geometry and spatial context of this domain while also revealing the limits of this specialization.
A. Rusnak, S. Kovalenko, Jingru Wang et al.· 0 citations
CVLNet is presented, a Cross-View Learning Network that predicts street-level perception from AlphaEarth embeddings and multi-source urban contextual data without requiring SVI at inference, enabling a more comprehensive assessment of urban environmental inequality.
Peilin Li, Pengfei Chen, Jing-Yu Wang et al.· 0 citations
Visual geo-localization seeks to enable autonomous aircraft positioning in GNSS-denied environments through large-scale image retrieval. Most existing methods, however, primarily rely on visible light images and are thus inherently sensitive to illumination changes and adverse weather conditions. Near-infrared (NIR) im...
Teng-Da Zhang, Yun-Zhou Zhang, Li Wang et al.· IEEE Transactions on Geoscie...· 0 citations
Location representations provide mobility models with fundamental information about the spatial position, functional characteristics, and relationships of places. However, existing embeddings are often dependent on mobility observations, unable to represent unseen locations, and weakly constrained to retain geographic...
Xing-Lei Wang, Stephen Law, Zi-Chao Zeng et al.· 0 citations
3D grounding aims to localize target entities in complex scenes from natural language and plays a fundamental role in embodied perception and spatial reasoning. However, existing approaches mostly rely on feature similarity or direct matching, making it difficult to connect natural-language intent with the implicit sem...
Shuai Zhang, Hong-Ye Hou, Qing-He Liu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.