STSFusion: segmentation task-driven spatial-frequency collaborative fusion for infrared and visible images
Abstract
Infrared and visible image fusion aims to simultaneously preserve the saliency of thermal targets and rich texture details. Most existing methods primarily rely on spatial-domain representations, while the frequency-domain information is not sufficiently explored. Moreover, it remains challenging to simultaneously maintain visual structures and enhance semantic discriminability. To address these issues, this paper proposes STSFusion, a segmentation task-driven spatial-frequency collaborative fusion method for infrared and visible images. Specifically, a shallow feature extraction encoder first maps infrared and visible images into a unified feature space. Then, a spatial-frequency collaborative feature decomposition module is constructed to model global basic information and local detail information in the spatial domain. Meanwhile, an adaptive frequency-domain guidance module introduces low-frequency and high-frequency responses to guide spatial global structure modeling and local detail representation, respectively. Furthermore, a differential fusion module is designed to generate adaptive gating weights according to the explicit differences between global and local features, thereby dynamically adjusting the information preservation ratio between them. Finally, a deep semantic-guided decoder is introduced to guide fused image reconstruction through semantic modulation. To further enhance the semantic discriminability of fused features, a segmentation task loss is introduced during training, enabling semantic, region, and boundary supervision to jointly optimize the network. Experimental results on multiple public datasets demonstrate that the proposed method achieves superior fusion performance over state-of-the-art methods and shows favorable task adaptability in downstream semantic segmentation and object detection.