Aug 2026· IEEE Transactions on Pattern Analysis and Machine Intelligence· Vol PP, pp. 1-18· 1 citation
Medicine
TL;DR
This paper proposes a symbiotic evolutionary learning framework for task-adaptive IVIF, termed EvoFuse, which enables a mutual adaptation process where the fusion network and task models are jointly updated under a Pareto-inspired non-degradation criterion.
Abstract
Infrared and visible image fusion (IVIF) targets to integrate thermal saliency and rich textures into a single image that is not only visually appealing but also beneficial to downstream vision tasks. However, conventional methods relying on heuristic visual criteria struggle to guarantee task utility. Conversely, task-driven fusion paradigms typically employ fixed-weight scalarization, which suffers from potentially conflicting objectives among heterogeneous tasks, leading to rigid compromises and sub-optimal generalization in multi-task scenarios. To overcome these bottlenecks, this paper proposes a symbiotic evolutionary learning framework for task-adaptive IVIF, termed EvoFuse. Rather than relying on static loss weights, we formulate the fusion-perception correlation from a multi-objective perspective. To structurally instantiate this formulation, we first develop a re-parameterizable fusion architecture that accommodates multi-branch representational capacity during training, yet analytically folds into an ultra-compact single-branch model for efficient inference. To navigate the conflicting multi-task objectives, we introduce an evolutionary search mechanism that dynamically evolves task-aware loss-weight configurations. This enables a mutual adaptation process where the fusion network and task models are jointly updated under a Pareto-inspired non-degradation criterion. Furthermore, a novel saliency discriminative loss is designed to explicitly emphasize semantically crucial regions. Extensive experiments across eleven datasets covering fusion and downstream perception tasks demonstrate that the proposed method achieves competitive or better results in most evaluated metrics, while maintaining efficient inference under the considered task settings.
Infrared-visible image fusion (IVIF) is pivotal for multimodal perception, yet reconciling the inherent information disparity between thermal and textural features remains a fundamental challenge. Existing prior-guided methods often rely on static constraints that induce optimization conflicts or utilize extrinsic sema...
Yi Shi, Huichao Xie, Yuqing Wang et al.· 0 citations
Infrared and visible image fusion is pivotal for robust visual perception across all weather conditions and scenes. Although deep learning-based methods have made notable progress, most either assume pre-aligned inputs or rely on implicit feature-space alignment, which fails to fundamentally address the amplification o...
Jin-Yuan Liu, Zengxi Zhang, Jiahao Zhang et al.· IEEE Transactions on Pattern...· 1 citation
Experimental results demonstrate that adaptability is not solely determined by model size, but rather by how effectively parameter plasticity is regulated in dynamic environments.
Xiao-Rong Zeng, Weiqiang Chen, Peng Shi et al.· 0 citations
Vision-language models exhibit remarkable zero-shot capabilities but suffer significant performance degradation under distribution shifts. While test-time adaptation (TTA) via Low-Rank Adaptation offers a parameter-efficient solution, we identify a fundamental bottleneck in current methods: the reliance on static rank...
Gengyuan Liu, Nan Wang, Chang Liu et al.· 0 citations
Generalist multitasking vision models aim to unify multiple vision tasks within a single framework, enabling more efficient and versatile learning. However, handling diverse vision tasks -- spanning dense and sparse predictions -- remains challenging due to their inherently varying output structures. In this paper, we...
Mohammad Mahdi, Nedyalko Prisadnikov, Yu-Qian Fu et al.· 0 citations
This work proposes MSGAN, a multi-scale global-local collaborative learning framework that integrates multi-scale mixed convolution and adaptive global-local attention to enhance feature representation and advances robust SOD for complex real-world scenarios and provides insights into attention-guided visual perception...