Modality-Aware Fusion for Optical and SAR Images via Complementary Structure-Texture Decomposition and Enhancement
Abstract
Fusion of optical and synthetic aperture radar (SAR) images effectively integrates their complementary information, enhancing the robustness of remote sensing interpretation. However, existing methods often struggle with insufficient adaptation to modality-specific characteristics, uncoordinated fusion of structure and texture, and limited detail representation, leading to spectral distortion, artifacts, and poor perceptual hierarchy. To address these challenges, this article proposes a novel modality-aware fusion method for optical and SAR images (MAOSF). The core of MAOSF is a complementary structure-texture decomposition (CSTD) and enhancement framework. First, a CSTD strategy integrates multiorientation relative total variation and rolling guided filter to decompose source images into structure and texture components in a modality-aware manner. Then, for the decomposed components, a cross-fusion with hierarchical enhancement is designed: the saliency-guided cross-structure fusion strategy preserves geometric contours and salient edges, while the dominant direction response-based cross-texture fusion strategy enhances fine details and suppresses speckle noise. Furthermore, a bitplane detail enhancement method is proposed to refine structural information in low-contrast regions using high-order bitplanes, significantly improving local detail visibility and visual hierarchy. Extensive experiments on three public datasets demonstrate that MAOSF achieves superior performance in structural fidelity, texture representation, and overall perceptual quality, outperforming several state-of-the-art fusion methods. Quantitatively, MAOSF outperforms the second-best methods by average margins of 10.9% in visual information fidelity, 7.7% in average gradient, and 9.6% in Qabf across the three benchmark datasets, confirming its consistent and substantial superiority.