A context-aware referring expression segmentation model for remote sensing that provides a solution for accurate multi-scale target segmentation guided by natural language descriptions in complex remote sensing scenes.
Abstract
Remote sensing referring image segmentation faces critical challenges including arbitrary target rotation, drastic scale variation, cluttered complex backgrounds, and large visual-language semantic gaps. Existing mainstream segmentation models adopt fixed-receptive-field backbones, coarse unidirectional cross-modal interaction and static learnable object queries, which easily cause small-object missed detection, blurred boundary segmentation and fragmented predictions. This paper proposes a context-aware referring expression segmentation model for remote sensing to tackle the above limitations. We construct a dual-stream feature extraction backbone with InternImage and CLIP text encoder, and design a semantic prior localization map module to generate spatial heatmaps for spatial inductive bias and improve small-object localization recall. A cross-modal context aggregator performs multi-scale bidirectional visual-text alignment, while a dynamic query initialization strategy and a language-guided Transformer decoder progressively improve target localization and mask refinement. Experiments are conducted on two standard RRSIS benchmarks, RefSegRS and RRSIS-D. On RefSegRS, the proposed model achieves an mIoU of 70.43% and an oIoU of 77.34%, and obtains significant improvements on small vehicles, buildings and slender road markings. On the larger-scale RRSIS-D dataset, our model achieves an mIoU of 63.71% and an oIoU of 73.68%. The proposed model provides a solution for accurate multi-scale target segmentation guided by natural language descriptions in complex remote sensing scenes.
DLPANet is proposed, a novel dual-level prototype alignment network centered on Prototype-Guided Spatial Attention, enabling simultaneous modeling of scene context and fine-grained details and demonstrates that the decoupled dual cross-attention mechanism provides superior prototype-query alignment compared to prior gl...
Mustafa Alawadi, M. Fateh· Jordanian Journal of Compute...· 0 citations
Rapid advancements in vision-language models have propelled Referring Remote Sensing Image Segmentation (RRSIS) to the forefront of Earth observation. However, practical deployments suffer severe performance degradation under a coupled dual-drift paradigm: visual domain drift from cross-spatial-resolution mismatches an...
Quan-Wei Liu, Tao Huang, Jia-Qi Yang et al.· 0 citations
Referring remote sensing image segmentation (RRSIS) aims to segment target ground objects in high-resolution remote sensing imagery according to textual descriptions. Due to complex backgrounds and large-scale variations, accurately aligning linguistic semantics with spatial regions remains a challenging problem. Exist...
Sen Lei, Shuai Li, Xin-Yu Xiao et al.· IEEE Transactions on Geoscie...· 2 citations
Open-vocabulary semantic segmentation (OVSS) of remote sensing faces severe performance degradation when encountering unseen scene distributions caused by geographic, sensor, and resolution variations. Existing vision–language approaches provide strong semantic priors but lack scene-invariant structural representations...
Wu-Biao Huang, Hu-Chen Li, Shuai Zhang et al.· IEEE Transactions on Geoscie...· 0 citations
Weakly supervised remote sensing semantic segmentation aims to achieve pixel-level prediction with limited annotation costs. Recently, CLIP-based methods have shown promising potential by leveraging vision-language alignment for semantic localization; however, they often suffer from ambiguous semantic representations a...
Bei Cheng, Quan-Li Deng, Tao Shen et al.· IEEE Geoscience and Remote S...· 0 citations
Open-vocabulary semantic segmentation (OVS) of remote sensing imagery is a challenging pixel-level task requiring strong generalization and adaptation to the spatial characteristics of remote sensing data. Although existing vision-language foundation models perform well in general domains, their image-level classificat...
Chang-Hao Zhao, Ling-Lin Zeng, Hai Liu· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.