Skip to content

M3RIS: Mutual Modulation Mamba Network for Referring Remote Sensing Image Segmentation

2026 · IEEE Transactions on Geoscience and Remote Sensing · Vol 64, pp. 5640116-5640116 · 0 citations · 70 references

Abstract

Referring remote sensing image segmentation (RRSIS) aims to map natural language queries to pixel-level masks within complex geographic scenes, requiring precise cross-modal semantic alignment. In recent years, Mamba has demonstrated impressive performance in various remote sensing tasks due to its powerful contextual modeling capability, offering new potential for RRSIS. However, vanilla Mamba primarily models sequential dependencies and provides limited support for cross-modal interaction, thus hindering precise vision–language semantic alignment. To overcome this limitation, we propose M3referring image segmentation (RIS), a mutual modulation Mamba network for RRSIS. Specifically, we propose the vision–language mutual modulation (VLMM) module that distills semantic knowledge (SK) from the segment anything model (SAM) to provide auxiliary semantic cues for cross-modal interaction. The module not only uses the distilled knowledge to guide the prediction of channel-wise affine coefficients and spatial dynamic convolution kernels to enhance target-relevant visual responses, but also introduces this knowledge into cross-attention as a semantic bias to enable language features to aggregate semantically relevant visual context. This facilitates mutual enhancement between vision and language, promoting more effective cross-modal alignment than conventional unidirectional modulation. Moreover, we introduce the hybrid semantic refined (HSR) decoder, which alternates between visual state space (VSS) blocks and cross-modal semantic refinement (CSR) modules. The VSS blocks capture global semantics via long-range context modeling, while the CSR modules derive multiscale dynamic convolution kernels from language features to refine local visual features through dynamic filtering, thereby enhancing semantic awareness of the referred object. Extensive experiments on RRSIS-D, RISBench, and RefSegRS datasets demonstrate that M3RIS outperforms several previous state-of-the-art approaches.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.