Skip to content
Open access

CLIP-RSICD: A Fine-Tuned Vision Transformer for Urban Land Cover Classification

Aug 2026 · Semina: Ciências Exatas e Tecnológicas · 0 citations · 52 references

Abstract

Reliable mapping of intra-urban land-cover categories at fine spatial and thematic resolutions is fundamental to evidence-based spatial planning. Nevertheless, transferring classification models across territories remains difficult because urban morphology varies regionally. This challenge is pronounced in South American urban centers, where rapid expansion alters metropolitan dynamics, demanding finegrained land monitoring to support zoning policies, infrastructure distribution, and ecological protection. Addressing this issue, we evaluate a remote-sensing (RS) adapted Contrastive Language–Image Pre-Training (CLIP) visual encoder for multi-class urban land cover (ULC) categorization. The proposed taxonomy incorporates ten classes, spanning high-, medium-, and low-density built environments, industrial complexes, vegetative covers, and exposed soil. The high-resolution (HR) ULC products offer a rigorous empirical framework for urban morphological assessment and planning applications in intermediate Latin American cities. Using HR, open-access satellite imagery, the framework was implemented across the central urban cluster of the Londrina Metropolitan Region, Southern Brazil. The domain-adapted CLIP model reached an overall accuracy of 93%, outperforming nine baseline deep learning architectures. Our findings highlight the capacity of vision-language backbones to capture spatial patterns and extract discriminative representations from RS data, providing a scalable solution for urban spatial analysis.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.