Cross-Modal Image Fusion via Structure Preservation and Detail Enhancement Optimization
Abstract
Highlights What are the main findings? Proposes Structure–Detail Constrained Fusion (SDC-Fusion), a frequency-decoupled infrared and visible image fusion framework that separately models low-frequency structure and high-frequency detail. Restricts Rectified Flow to the high-frequency wavelet subbands and learns a gated few-step residual compensation trajectory from the visible high-frequency component to a local-directional-energy-guided target, rather than generating the complete fused image or latent representation. Develops a low-frequency structure preservation module that integrates a four-directional Mamba scan with multi-dilated depthwise convolutions, simultaneously maintaining global luminance integrity and optimizing local grayscale transitions. Ranks first on five of seven metrics and second on the remaining two metrics on both the MSRS and M3FD datasets, while using 0.535 M parameters and 67.5 G FLOPs. Abstract Existing end-to-end infrared–visible fusion methods often blur edges, smooth textures and weaken target-to-background contrast. We therefore propose Structure–Detail Constrained Fusion (SDC-Fusion), a frequency-decoupled framework with separate constraints on structure and detail. The proposed method employs the Haar wavelet transform to decompose the source images into low-frequency structural and high-frequency detail components. The high-frequency branch uses a local directional-energy prior to construct the target guidance. Unlike existing Rectified Flow-based approaches that operate on the full image or a generic latent representation, our gated module applies Rectified Flow only to the high-frequency wavelet subbands. It learns a few-step residual trajectory from the visible high-frequency coefficients to the target representation, enhancing infrared target boundaries and visible textures without altering low-frequency structure. In the low-frequency branch, adaptive weighting, four-directional Mamba scanning, and multi-dilation depthwise convolutions are integrated to preserve global luminance and background structure while optimizing local grayscale transitions. Comparative experiments against eleven representative fusion methods on the MSRS and M3FD datasets show that SDC-Fusion ranks first in SSIM, VIF, Qabf, SF, and PSNR on MSRS, and first in SSIM, VIF, Qabf, SD, and PSNR on M3FD, while ranking second in the remaining two metrics on each dataset. Relative to the strongest competing result, the largest improvements reach 10.44% in SF on MSRS and 5.72% in VIF on M3FD. The model contains 0.535 M parameters and requires 67.5 G FLOPs.