Skip to content

MMF-Net: Multimodal Mutual Feedback Fusion Network for Referring Remote Sensing Image Segmentation

2026 · IEEE Geoscience and Remote Sensing Letters · Vol 23, pp. 2506705-2506705 · 0 citations · 22 references

Abstract

Referring remote sensing image segmentation (RRSIS) commonly treats language as a fixed query that only modulates visual features. This open-loop design is brittle in aerial scenes containing repeated objects, weak appearance cues, and relational expressions: visual evidence cannot revise which words and relations should dominate the query. We present MMF-Net, a closed-loop fusion architecture that couples stagewise reciprocal token refinement with decoder-wide context propagation. Its vision–language mutual feedback fusion (VLMFF) module updates visual and linguistic tokens through symmetric cross-attention and gated residual fusion, while multimodal context-aware fusion (MCAF) consolidates the corefined state and broadcasts it across the visual hierarchy. Under an identical Swin-B+ Bidirectional Encoder Representations from Transformer (BERT) backbone, decoder, training schedule, and data split, MMF-Net improves a strong internal baseline by 2.87/2.51/1.48 points at Pr@0.5/0.7/0.9 and by 0.68 oIoU and 0.58 mIoU. A factorial ablation further exposes positive VLMFF–MCAF interaction effects of 5.54 points at Pr@0.7 and 2.40 mIoU. MMF-Net reaches 64.86% mIoU and 78.14% oIoU on RRSIS-D with 115.2M parameters and 68.7 GFLOPs.

View source

Similar papers

2026

M3RIS: Mutual Modulation Mamba Network for Referring Remote Sensing Image Segmentation

Referring remote sensing image segmentation (RRSIS) aims to map natural language queries to pixel-level masks within complex geographic scenes, requiring precise cross-modal semantic alignment. In recent years, Mamba has demonstrated impressive performance in various remote sensing tasks due to its powerful contextual...

Hu Guo, Bin Sun, Wei-Qing Lu et al. · 0 citations
Preprint Sep 2026

VPRef: A Cross-Domain Benchmark for Referring Remote Sensing Image Segmentation

Rapid advancements in vision-language models have propelled Referring Remote Sensing Image Segmentation (RRSIS) to the forefront of Earth observation. However, practical deployments suffer severe performance degradation under a coupled dual-drift paradigm: visual domain drift from cross-spatial-resolution mismatches an...

Quan-Wei Liu, Tao Huang, Jia-Qi Yang et al. · 0 citations
Sep 2026

MCMB-UNet: A dual-encoder network with multi-attention for remote sensing image segmentation

MCMB-UNet is proposed, an effective dual-path encoder model with multi-attention mechanisms that achieves a favorable accuracy-efficiency trade-off compared to mainstream Transformer-based models, and demonstrates good applicability on the Inria Aerial Labeling and Massachusetts Buildings datasets.

Dong-Dong Huang, Yu-Hong Ding · 0 citations
Open access Aug 2026

MTC-Net: Leveraging Multi-Temporal Consistency and Multi-View Synergistic Contrastive Learning for Remote Sensing Scene Classification

A progressive layer-wise contrastive learning framework (MTC-Net) that couples the pseudo-label with the network’s representational hierarchy, forming a curriculum from local texture robustness to global semantic invariance.

Xiao Xiao, Han Zhang, Kenan Cheng et al. · 0 citations
Open access 2026

Dual-Level Prototype Alignment via Cross-Attention for Few-Shot Remote Sensing Semantic Segmentation

DLPANet is proposed, a novel dual-level prototype alignment network centered on Prototype-Guided Spatial Attention, enabling simultaneous modeling of scene context and fine-grained details and demonstrates that the decoupled dual cross-attention mechanism provides superior prototype-query alignment compared to prior gl...

Mustafa Alawadi, M. Fateh · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.