Skip to content

Edge-guided attention and multi-scale fusion for high-quality multi-modal image fusion

Sep 2026 · Multimedia Systems · Vol 32 · 0 citations · 35 references

TL;DR

This work proposes EdgeAttnSwin, a framework that reformulates fusion as a three-stage progressive optimization process with explicit causal dependencies, and demonstrates that this method outperforms state-of-the-art approaches on infrared-visible and medical image fusion benchmarks.

View source

Similar papers

Open access Sep 2026

HDPF: Hierarchical Dual-Perspective Collaborative Modeling for Multimodal Image Fusion

Different imaging modalities exhibit inherent discrepancies in intensity characteristics and information representation, resulting in pronounced heterogeneity among multimodal features. When these heterogeneous features are directly learned and fused within a unified representation space, feature coupling may arise, le...

Zhi-Xiang Zhang, Qiang Tang, Zhong Zhang et al. · 0 citations
Open access Sep 2026

Boundary-Guided Dual-Perspective Cross-Modal Fusion Network for RGB-IR Object Detection

A Boundary-Guided Dual-Perspective Cross-Modal Fusion Network (BDPNet) is proposed to explicitly preserve shallow geometric structures and decouple deep semantic fusion into macroscopic and microscopic perspectives.

Hu Lin, Zhi-Wei Fu, Xiu-Mei Chen et al. · 0 citations
#edge computing Open access Sep 2026

Edge-Aware Dynamic Convolution and Cross-Modal Relation-Aware Fusion Network for Infrared and Visible Image Fusion

Infrared and visible image fusion combines complementary thermal and structural information from the two modalities into a single composite image. Existing methods have two critical limitations: (1) inadequate utilization of visible structural information causes blurred edges, and (2) modality-specific and shared respo...

Shun-Li Liu, An-Jie Chen, Qiao Luo et al. · 0 citations
Open access Aug 2026

SACMFuse: Structure-Aware Cross-Modal Interaction Network for Multi-Modal Image Fusion

Multi-modal image fusion integrates complementary information from different modalities to generate a unified representation that is informative for human perception and beneficial to downstream vision tasks. However, existing methods often inefficiently model global features and their cross-modal interaction is insuff...

Quan-Rui Wen, Xiu Shu, Xinming Zhang et al. · 0 citations
Preprint Sep 2026

When Integral Meets Decomposition: A Signal-Level Self-Supervised Feature Decompose Paradigm for Multi-Modal Image Fusion

Multimodal image fusion (MMIF) aims to integrate complementary information from different modalities into a high-quality fused image and support downstream tasks. Recently, feature decomposition has become an important paradigm by separating source images into common and modality-specific unique features. However, exis...

Ze-Yu Wang, Jia-Yu Wang, Hai-Yu Song et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.