FastSAM-Guided Adversarial Learning for Unsupervised Multimodal Remote Sensing Change Detection
Abstract
Multimodal change detection (CD), due to its ability to flexibly adapt to data acquired from different types of sensors, has become an important research direction in the field of remote sensing. However, existing methods generally lack feature representations with sufficient generalization capacity, leading to pronounced performance degradation when applied across diverse scenes. To address this limitation, we propose a novel FastSAM-guided adversarial learning (FastSAM-GAL) framework for unsupervised multimodal CD. The proposed framework effectively mitigates cross-modal discrepancies via adversarial learning and constructs a dual-branch architecture comprising a FastSAM encoder and a remote sensing encoder, thereby fully exploiting the complementary information between generic visual semantic features and domain-specific remote sensing representations to substantially enhance cross-scene generalization capability. Furthermore, a progressive multiscale perception adaptor (PMSPA) and a spatial-frequency collaborative edge augmentation (SFCEA) module are specifically designed to further improve FastSAM-GAL’s ability to perceive objects at different scales and capture change boundaries. Finally, comparative experiments on five real-world multimodal CD datasets against 12 state-of-the-art methods demonstrate that the proposed FastSAM-GAL maintains stable and reliable detection performance across multimodal remote sensing data with diverse scenes, with its $\kappa $ metric significantly outperforming all competing methods.