Skip to content
Preprint

Cross-View Urban Sensing: Mapping Subjective Streetscape Perception via AlphaEarth Embeddings and Urban Context

Aug 2026 · 0 citations · 46 references
Computer Science

TL;DR

CVLNet is presented, a Cross-View Learning Network that predicts street-level perception from AlphaEarth embeddings and multi-source urban contextual data without requiring SVI at inference, enabling a more comprehensive assessment of urban environmental inequality.

Abstract

Residents'perception of the urban streetscape is an important factor in public health, active mobility, and social wellbeing. Street view imagery (SVI) has emerged as a widely used data source for assessing these perceptual qualities, yet its uneven coverage and irregular updating limit large-scale measurement. Here, we present CVLNet, a Cross-View Learning Network that predicts street-level perception from AlphaEarth embeddings and multi-source urban contextual data without requiring SVI at inference. CVLNet applies per-task adaptive gating to jointly model five perceptual dimensions, using labels from the pretrained SVI-Percept model as ground truth. The proposed method is evaluated across four Southeast Asian cities: Singapore, Kuala Lumpur, Jakarta, and Manila. CVLNet achieves a median road-segment-level Adjusted $R^{2}$ of 0.76 and consistently outperforms the baseline models, with gains ranging from 5.9--11.3% across the five perceptual dimensions. Ablation experiments show that AlphaEarth features and urban contextual features contribute complementary information. We further produce citywide road-level streetscape perception maps for five subjective perceptual dimensions across all four cities, extending perception estimation from the 13--31% of the road network directly covered by available SVI to the complete road network of each city. Integrating these maps with WorldPop gridded population data, we quantify exposure inequality across population-density, demographic, and land-use groups using the Deficit Palma Ratio. These results demonstrate that remote sensing can serve as a scalable alternative to SVI for citywide streetscape perception mapping, enabling a more comprehensive assessment of urban environmental inequality.

View source

Similar papers

Review Aug 2026

From Street View Imagery to Street Quality Indicators: Vision Language Inference for the Suburban 15-minute City

This paper presents a planning-oriented assessment of streetscape qualities in the north-eastern periphery of Nice using the latest release of SAGAI (Streetscape Analysis with Generative AI), an open-source workflow that leverages vision-language models for large-scale streetscape analysis from Google Street View image...

Joan Perez, Giovanni Fusco · 0 citations
Review Open access Sep 2026

A Multidimensional Framework for Diagnosing Streetscape Perception in Historic-District Renewal Using Street-View Imagery and Deep Learning

The renewal of historic districts needs to address human-centered issues, such as cultural expression, spatial experience, and visual comfort, at the street scale. However, existing assessment methods predominantly rely on field surveys, expert judgment, or individual visual indicators, making it difficult to produce r...

Ji-Long Li, Hong-Yang Chen, Pan Liao et al. · 0 citations
Preprint Aug 2026

Vision Models Predict Urban Scene Appraisal with Limited Neural Alignment

Pretrained vision embeddings are increasingly used as general-purpose representations for modelling how people appraise urban scenes, and are validated almost entirely by how well they predict human ratings. High predictive accuracy does not establish that these embeddings organise scenes as human perception does. We t...

Kai-Zhen Tan, Yuan-Tao Deng · 0 citations
Book Open access Aug 2026

UrbanGraphEmbeddings: Learning and Evaluating Spatially Grounded Multimodal Embeddings for Urban Environments

This work proposes UGE, a two-stage training strategy that progressively and stably aligns images, text, and spatial structures by combining instruction-guided contrastive learning with graph-based spatial encoding, and introduces \dataset, a spatially grounded dataset that anchors street-view images to structured spat...

Jie Zhang, Xing-Tong Yu, Yuan Fang et al. · 0 citations
#artificial intelligence Review Aug 2026

Polis: 3D Self-Supervision at City Scale

The results show that distributionally-regularized joint embedding architectures can be successful on challenging city-scale 3D scenes, and that transfer improves when self-supervision is designed for the capture geometry and spatial context of this domain while also revealing the limits of this specialization.

A. Rusnak, S. Kovalenko, Jingru Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

When Correlations Mislead: Confounder-Aware Multi-View Urban Region Representation Learning

Urban region representation learning commonly combines heterogeneous data sources, such as mobility flows, points of interest, and land-use information, to support tasks including mobility analysis, public safety forecasting, and service demand estimation. Existing multi-view methods typically improve region embeddings...

Sean Bin Yang, Ying Sun, Zong-Yi Xu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.