Cross-View Change Detection via Self-Supervised Learning
This article addresses the problem of cross-view heterogeneous change detection using satellite images and autonomous aerial vehicle imagery, a setting characterized by severe viewpoint differences, scale variations, and sensing modality discrepancies, as well as the absence of reliable labels. To overcome these challenges, we propose a self-supervised contrastive and predictive learning framework that learns discriminative representations directly from unlabeled cross-view image pairs. The framework jointly exploits patch-level and pixel-level objectives to capture both fine-grained local changes and higher-level semantic consistency across views. In addition, a CutMix-based data augmentation strategy is introduced to improve representation diversity. We further enhance Swin Transformer with cross-attention and channel attention mechanisms to facilitate effective multiscale and cross-modal feature interaction. Experimental results on cross-view and bitemporal change detection datasets demonstrate that the proposed approach achieves robust and competitive performance without relying on precise geometric alignment or manual annotations, highlighting its practical relevance for real-world Earth observation applications involving heterogeneous and rapidly acquired data.