SACMFuse: Structure-Aware Cross-Modal Interaction Network for Multi-Modal Image Fusion
Abstract
Multi-modal image fusion integrates complementary information from different modalities to generate a unified representation that is informative for human perception and beneficial to downstream vision tasks. However, existing methods often inefficiently model global features and their cross-modal interaction is insufficient, resulting in the suboptimal preservation of structural details and modality-specific information. To address these issues, we propose SACMFuse, a structure-aware cross-modal interaction network for multi-modal image fusion. SACMFuse is built upon the linear complexity attention mechanism of LAMA, which enables efficient and effective global feature modeling. In this study, LAMA is extended into a CrossLAMA mechanism to facilitate deep-level information interaction among modalities. Within this framework, we propose a Frequency–Spatial Rectification Module (FSRM) that jointly models spatial-domain structures and frequency-domain representations in a unified manner. By perceiving and adaptively rectifying structural features across modalities, FSRM enhances the structural consistency and discriminability of the fused features. Furthermore, to strengthen cross-modal complementarity, we design a Difference-driven Cross-modal Interaction Module (DCIM), where inter-modal discrepancies explicitly guide information exchange among modalities. This mechanism encourages the network to focus on complementary structures while suppressing redundant responses. Extensive experiments on multiple datasets demonstrate that the performance of SACMFuse is competitive for the majority of metrics, providing a robust and advanced solution for image fusion tasks.