A Hierarchical Structure-Guided Cross-Modal Calibration Network for SAR–Optical Cloud Removal
Abstract
Cloud contamination remains a fundamental obstacle for optical Earth observation, severely limiting the availability and temporal continuity of remote sensing imagery. Although synthetic aperture radar (SAR) provides complementary all-weather structural information, existing SAR–optical cloud-removal methods remain challenged by three coupled issues: heterogeneous microwave-scattering and optical-reflectance characteristics complicate cross-modal feature integration; cloud occlusions vary substantially in spatial extent and therefore require adaptive contextual support; and reconstruction must preserve global scene consistency and local spectral–textural details without introducing excessive computational redundancy. To address these challenges, we propose a hierarchical structure-guided cross-modal calibration network (HySADNet) for SAR-assisted optical cloud removal. HySADNet adopts an asymmetric dual-stream architecture in which the optical pathway serves as the primary reconstruction stream, while the SAR pathway provides cloud-independent structural guidance. A cross-modal synergistic calibration and fusion module performs reciprocal feature recalibration and reliability-weighted fusion between the two modalities, while a dynamic kernel fusion (DKF) module adaptively aggregates multiscale contextual information. In addition, a structure-aware hybrid Transformer block (SAHTB) combines window-based self-attention with convolutional local priors to jointly model long-range dependencies and fine-grained structures. Experiments on SEN12MS-CR and LuojiaSET-OSFCR demonstrate the effectiveness of the proposed method. Compared with the corresponding second-best results, HySADNet improves peak signal-to-noise ratio (PSNR) by 0.2608 dB on SEN12MS-CR and by 0.7050 dB on LuojiaSET-OSFCR; it also reduces spectral angle mapper (SAM) by 0.1612 and 0.2947, respectively. Downstream semantic-segmentation evaluation further indicates that the reconstructed imagery preserves useful structural and spectral information for the evaluated high-level vision task.