2026· IEEE Geoscience and Remote Sensing Letters· Vol 23, pp. 2506705-2506705· 0 citations· 22 references
Abstract
Referring remote sensing image segmentation (RRSIS) commonly treats language as a fixed query that only modulates visual features. This open-loop design is brittle in aerial scenes containing repeated objects, weak appearance cues, and relational expressions: visual evidence cannot revise which words and relations should dominate the query. We present MMF-Net, a closed-loop fusion architecture that couples stagewise reciprocal token refinement with decoder-wide context propagation. Its vision–language mutual feedback fusion (VLMFF) module updates visual and linguistic tokens through symmetric cross-attention and gated residual fusion, while multimodal context-aware fusion (MCAF) consolidates the corefined state and broadcasts it across the visual hierarchy. Under an identical Swin-B+ Bidirectional Encoder Representations from Transformer (BERT) backbone, decoder, training schedule, and data split, MMF-Net improves a strong internal baseline by 2.87/2.51/1.48 points at Pr@0.5/0.7/0.9 and by 0.68 oIoU and 0.58 mIoU. A factorial ablation further exposes positive VLMFF–MCAF interaction effects of 5.54 points at Pr@0.7 and 2.40 mIoU. MMF-Net reaches 64.86% mIoU and 78.14% oIoU on RRSIS-D with 115.2M parameters and 68.7 GFLOPs.
Referring remote sensing image segmentation (RRSIS) aims to map natural language queries to pixel-level masks within complex geographic scenes, requiring precise cross-modal semantic alignment. In recent years, Mamba has demonstrated impressive performance in various remote sensing tasks due to its powerful contextual...
Hu Guo, Bin Sun, Wei-Qing Lu et al.· IEEE Transactions on Geoscie...· 0 citations
Rapid advancements in vision-language models have propelled Referring Remote Sensing Image Segmentation (RRSIS) to the forefront of Earth observation. However, practical deployments suffer severe performance degradation under a coupled dual-drift paradigm: visual domain drift from cross-spatial-resolution mismatches an...
Quan-Wei Liu, Tao Huang, Jia-Qi Yang et al.· 0 citations
MCMB-UNet is proposed, an effective dual-path encoder model with multi-attention mechanisms that achieves a favorable accuracy-efficiency trade-off compared to mainstream Transformer-based models, and demonstrates good applicability on the Inria Aerial Labeling and Massachusetts Buildings datasets.
Dong-Dong Huang, Yu-Hong Ding· Signal, Image and Video Proc...· 0 citations
A context-aware referring expression segmentation model for remote sensing that provides a solution for accurate multi-scale target segmentation guided by natural language descriptions in complex remote sensing scenes.
A progressive layer-wise contrastive learning framework (MTC-Net) that couples the pseudo-label with the network’s representational hierarchy, forming a curriculum from local texture robustness to global semantic invariance.
Xiao Xiao, Han Zhang, Kenan Cheng et al.· Remote Sensing· 0 citations
DLPANet is proposed, a novel dual-level prototype alignment network centered on Prototype-Guided Spatial Attention, enabling simultaneous modeling of scene context and fine-grained details and demonstrates that the decoupled dual cross-attention mechanism provides superior prototype-query alignment compared to prior gl...
Mustafa Alawadi, M. Fateh· Jordanian Journal of Compute...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.