Skip to content
Conference

Lightweight infrared and visible image fusion via global selective state-space modeling

Aug 2026 · International Conference on Digital Image Processing · Vol 14351, pp. 143510C - 143510C-11 · 0 citations · 25 references
Engineering

TL;DR

A knowledge distillation-based training strategy, together with a structural consistency constraint, is introduced to enhance the fusion quality of the lightweight model by guiding the student network to inherit discriminative representations from the teacher model while preserving a compact architecture.

Abstract

Infrared and visible image fusion aims to effectively exploit the strength of the infrared modality in target saliency perception and the complementary capability of the visible modality in representing fine-grained texture details, which is crucial for complex scene understanding and downstream visual tasks. Existing deep learning methods achieve strong performance but often rely on complex cross-modal modeling, leading to large models and high computational costs, limiting real-time deployment. To address these challenges, this paper proposes a lightweight infrared and visible image fusion framework, termed GSSMFuse, based on Global Selective State-Space Modeling. In the feature extraction stage, depthwise separable convolutions are employed to capture local spatial structural information with low computational overhead, followed by a Mamba-based GSSM block to efficiently model long-range dependencies, enabling joint representation of local details and global semantic information. In the fusion stage, infrared and visible features processed by the GSSM block are directly combined and further integrated using lightweight convolution, effectively exploiting the complementary characteristics of the two modalities while maintaining high computational efficiency. Furthermore, a knowledge distillation-based training strategy, together with a structural consistency constraint, is introduced to enhance the fusion quality of the lightweight model by guiding the student network to inherit discriminative representations from the teacher model while preserving a compact architecture. Extensive experiments on the public M3FD dataset demonstrate that the proposed GSSMFuse consistently outperforms existing state-of-the-art fusion methods, while significantly reducing model parameters and computational complexity and achieving competitive performance in downstream object detection tasks.

View source

Similar papers

Open access Aug 2026

Infrared and Visible Image Fusion via Style-Based Recalibration and Edge Enhancement

A lightweight end-to-end IVIF network with two complementary refinement modules that achieves the best or tied-best value on three of seven standard fusion-quality metrics on FMB and four of seven on LLVIP, and ablation results further confirm the complementary effects of MSG and DGM.

Wenhua Zhao, Lei Zhong · 0 citations
Open access Aug 2026

STSFusion: segmentation task-driven spatial-frequency collaborative fusion for infrared and visible images

Infrared and visible image fusion aims to simultaneously preserve the saliency of thermal targets and rich texture details. Most existing methods primarily rely on spatial-domain representations, while the frequency-domain information is not sufficiently explored. Moreover, it remains challenging to simultaneously main...

Yutong Chen, Zhe Hu, Yang Li et al. · 0 citations
Open access 2026

Infrared and Visible Image Fusion Based on Gaussian Weighted Standard Deviation Filter

Infrared and visible image fusion (IVIF) aims to generate a comprehensive and informative fused image by combining complementary thermal radiation and texture details from dual-modal source images. However, current fusion techniques still suffer from inadequate preservation of thermal targets and fine-grained details,...

Lian Liu, Xiao-L. Cheng, Jin-Liang Huang et al. · 0 citations
2026

Multiresolution Infrared and Visible Image Fusion via Implicit Neural Representations

Dual-spectral measurement instruments integrate complementary thermal radiation and textural information, thereby alleviating the information acquisition limitations of single-modality sensors. Due to detector manufacturing constraints and hardware cost, the infrared sensors usually provide lower spatial sampling rates...

Shu-Chen Sun, Li-Gen Shi, Jun Qiu et al. · 0 citations
Conference Sep 2026

Visible and infrared image fusion based on deep learning

Feature extraction is a common point at which current image fusion approaches fall short in establishing relationships between local and global information. The result is merged photos of low quality. This research presents an interactive transformer-based infrared and visible image fusion network to solve this problem...

Yu-Ming Wang, Kai-Xiang Liang, Rao Fu et al. · 0 citations
Open access Aug 2026

SFDNet: spatial-frequency decoupled network for infrared and visible image fusion

Infrared and visible image fusion aims to integrate complementary information from different modalities to generate images with both salient targets and rich textures. However, existing methods mainly rely on spatial feature modeling and lack an explicit mechanism to exploit frequency-aware representations, limiting th...

Yufeng Li, Lei Yu, Chuanlong Xie et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.