Skip to content
Book Open access

UrbanGraphEmbeddings: Learning and Evaluating Spatially Grounded Multimodal Embeddings for Urban Environments

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 10246-10257 · 0 citations · 65 references

TL;DR

This work proposes UGE, a two-stage training strategy that progressively and stably aligns images, text, and spatial structures by combining instruction-guided contrastive learning with graph-based spatial encoding, and introduces \dataset, a spatially grounded dataset that anchors street-view images to structured spatial graphs and provides graph-aligned supervision via spatial reasoning paths and spatial context captions.

Abstract

Learning transferable multimodal embeddings for urban environments is challenging because urban understanding is inherently spatial, yet existing datasets and benchmarks lack explicit alignment between street-view images and urban structure. We introduce \dataset, a spatially grounded dataset that anchors street-view images to structured spatial graphs and provides graph-aligned supervision via spatial reasoning paths and spatial context captions, exposing distance, directionality, connectivity, and neighborhood context beyond image content. Building on UGData, we propose UGE, a two-stage training strategy that progressively and stably aligns images, text, and spatial structures by combining instruction-guided contrastive learning with graph-based spatial encoding. Lastly, we introduce UGBench, a comprehensive benchmark to evaluate how spatially grounded embeddings support diverse urban understanding tasks, including geolocation ranking, image retrieval, urban perception, and spatial grounding. In particular, UGE is built on multiple state-of-the-art VLM backbones: Qwen2-VL, Qwen2.5-VL, Phi-3-Vision, and LLaVA1.6-Mistral, with fixed-dimensional spatial embeddings trained under LoRA tuning. UGE built upon Qwen2.5-VL-7B backbone achieves up to 44% improvement in image retrieval and 30% in geolocation ranking on training cities, and over 30% and 22% gains respectively on held-out cities, demonstrating the effectiveness of explicit spatial grounding for spatially intensive urban tasks. The code and datasets are available at https://github.com/Jaygagaga/UGE/tree/main.

Read PDF

Similar papers

Preprint Aug 2026

CoST: Semantic-Aware Urban Understanding via Spatial-Temporal Alignment

Geospatial representation learning from satellite imagery is a fundamental problem for large-scale urban analysis and real-world applications. Despite recent advances, current methods struggle with cross-region generalization and semantic interpretability due to their reliance on region-specific auxiliary data and the...

Yutian Jiang, Jiabo Liu, Xixuan Hao et al. · 1 citation
#artificial intelligence Review Aug 2026

Polis: 3D Self-Supervision at City Scale

The results show that distributionally-regularized joint embedding architectures can be successful on challenging city-scale 3D scenes, and that transfer improves when self-supervision is designed for the capture geometry and spatial context of this domain while also revealing the limits of this specialization.

A. Rusnak, S. Kovalenko, Jingru Wang et al. · 0 citations
Preprint Aug 2026

Cross-View Urban Sensing: Mapping Subjective Streetscape Perception via AlphaEarth Embeddings and Urban Context

CVLNet is presented, a Cross-View Learning Network that predicts street-level perception from AlphaEarth embeddings and multi-source urban contextual data without requiring SVI at inference, enabling a more comprehensive assessment of urban environmental inequality.

Peilin Li, Pengfei Chen, Jing-Yu Wang et al. · 0 citations
2026

Dual-Context Joint Representation for Near-Infrared Geo-Localization in Urban Environments

Visual geo-localization seeks to enable autonomous aircraft positioning in GNSS-denied environments through large-scale image retrieval. Most existing methods, however, primarily rely on visible light images and are thus inherently sensitive to illumination changes and adverse weather conditions. Near-infrared (NIR) im...

Teng-Da Zhang, Yun-Zhou Zhang, Li Wang et al. · 0 citations
#machine learning Preprint Aug 2026

LE4Mob: Towards Inductive, Distance-Aware and General-Purpose Location Embedding for Human Mobility Modelling

Location representations provide mobility models with fundamental information about the spatial position, functional characteristics, and relationships of places. However, existing embeddings are often dependent on mobility observations, unable to represent unseen locations, and weakly constrained to retain geographic...

Xing-Lei Wang, Stephen Law, Zi-Chao Zeng et al. · 0 citations
Preprint Sep 2026

CitySTAR: Structured and Topology-Aware Reasoning for Open-Vocabulary Urban 3D Grounding

3D grounding aims to localize target entities in complex scenes from natural language and plays a fundamental role in embodied perception and spatial reasoning. However, existing approaches mostly rely on feature similarity or direct matching, making it difficult to connect natural-language intent with the implicit sem...

Shuai Zhang, Hong-Ye Hou, Qing-He Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.