MpGI: Multiperspective Generation and Integration for High-Quality Remote Sensing Image–Text Datasets
Abstract
Recent advances in remote sensing (RS) vision-language foundation models (VLFMs) rely heavily on large-scale paired image–text data. However, existing dataset construction methods mainly emphasize data scale while overlooking a more fundamental limitation: the lack of structured and comprehensive semantic representation. Since RS imagery inherently contains multilevel and multiperspective semantics, including object attributes, spatial layouts, scene context, and functional relationships, single-pass caption generation often produces incomplete, biased, or weakly aligned descriptions. To address this issue, we propose a multiperspective semantic distillation paradigm for constructing high-quality RS image–text datasets. Based on this paradigm, we introduce a two-stage framework termed multiperspective generation and integration (MpGI). In the first stage, we perform annotation-guided knowledge elicitation using multiple multimodal large language models (MLLMs), where structured annotations are transformed into model-understandable instructions to extract diverse and complementary visual-semantic descriptions. In the second stage, we conduct semantic distillation and alignment-aware compression using large language models (LLMs), which integrate complementary descriptions and reorganize the resulting semantics into compact captions better suited for vision-language learning. With this framework, we construct HQRS-IT-210K, a high-quality RS image–text dataset containing approximately 210K images and 1.26M image–text pairs. We fine-tune CLIP to obtain HQRS-CLIP, which outperforms previous state-of-the-art RS vision-language models across multiple downstream tasks, including zero-shot classification (ZSC), few-shot classification (FSC), cross-modal retrieval, and semantic localization (SeLo), while using only 4.2% of the training data used by previous approaches. We further fine-tune CoCa on the full dataset to obtain RS-CoCa, which demonstrates strong RS image captioning (RSIC) capability and can generate captions comparable to manual annotations. These results show that improving semantic quality and image–text alignment is a more effective direction than merely increasing data volume for advancing RS vision-language modeling. The dataset, pretrained models, and code are available at https://github.com/YiguoHe/HQRS-210K-and-HQRS-CLIP