HDPF: Hierarchical Dual-Perspective Collaborative Modeling for Multimodal Image Fusion
Abstract
Different imaging modalities exhibit inherent discrepancies in intensity characteristics and information representation, resulting in pronounced heterogeneity among multimodal features. When these heterogeneous features are directly learned and fused within a unified representation space, feature coupling may arise, leading to mutual interference between global structural information and local fine-grained details. To address the limited differentiated modeling of structural and fine-grained information in existing methods, we propose Hierarchical Dual-Perspective Collaborative Modeling for Multimodal Image Fusion (HDPF). HDPF employs a Dual-Perspective Feature Aggregation Block (DFAB) to jointly exploit convolution-based local representation and Transformer-based global contextual modeling. Building upon this dual-perspective representation, HDPF further constructs two differentiated pathways dedicated to structural information and detail information, respectively, thereby providing differentiated representations of complementary multimodal information. The extracted features are subsequently reorganized and integrated by the decoder to reconstruct the final fused image. Extensive experiments are conducted on multiple infrared–visible image fusion (IVF) datasets as well as medical image fusion (MIF) tasks. On the TNO dataset, HDPF achieves VIF and MI scores of 0.80 and 3.50, respectively. On the MRI–PET fusion task, the VIF score reaches 0.61. The experimental results demonstrate that HDPF achieves competitive and relatively balanced fusion performance across the evaluated multimodal image fusion tasks.