It is argued that VI-ReID should be treated as an early cross-modal correspondence learning problem rather than only a late embedding alignment problem, and CMIA-Net is proposed, a framework that establishes bidirectional visible-infrared interaction at shallow backbone stages and introduces Spectral-Invariant Augmentation to stabilize early interaction.
Abstract
Visible-infrared person re-identification (VI-ReID) remains challenging due to severe spectral discrepancy and local cross-modal misalignment between visible and infrared images. Most existing methods alleviate this discrepancy through middle- or late-stage feature alignment, but modality-specific shallow features may have already accumulated spectral bias and local correspondence errors before shared representations are formed. In this paper, we argue that VI-ReID should be treated as an early cross-modal correspondence learning problem rather than only a late embedding alignment problem. To this end, we propose CMIA-Net, a framework that establishes bidirectional visible-infrared interaction at shallow backbone stages. Its core module, cross-modal interaction attention (CMIA), enables visible and infrared feature maps to exchange complementary local information before deep semantic aggregation, thereby reducing progressive stream divergence. To stabilize early interaction, we further introduce spectral-invariant augmentation (MC-Aug) to suppress over-reliance on visible-spectrum cues and a bi-directional hetero-center learning (BHCT) loss to improve class-center-level cross-modal compactness and inter-class separability. Experiments on RegDB and SYSU-MM01 show that CMIA-Net achieves strong performance on RegDB and the SYSU-MM01 indoor-search protocol, while remaining competitive under the more challenging SYSU-MM01 all-search setting. Specifically, CMIA-Net obtains 94.09% Rank-1 accuracy and 90.74% mAP in the RegDB visible-to-thermal setting, and 66.67% Rank-1 accuracy and 58.12% mAP in the SYSU-MM01 all-search setting.
Experiments show that DIGCA achieves competitive overall performance compared with recent VI-ReID methods, providing empirical support for the effectiveness of the proposed decoupling-guided alignment strategy in cross-modal identity matching.
This study proposes a novel dual-path convolution based multi-scale feature alignment (DCMFA) network that significantly outperforms existing mainstream methods in terms of recognition accuracy.
Bailiang Huang, Bin Chen, Tian-Ran Sun et al.· 電腦學刊· 0 citations
A Progressively Biased Split Vision Transformer (PBSVT) is proposed, which combines a split ViT backbone with progressive bias training to gradually reduce RGB-dominant bias while preserving modality-shared structure and demonstrates the effectiveness of progressive modality transition for robust VI-ReID representation...
Mengru Jiao, Xin-Yue Xu, Jun-Feng Zhang· International journal of pat...· 0 citations
A novel Robust Modality Unified Learning framework (RMUL), consisting of a Robust Cross-modality Label Transfer (RCLT) method and Modality Unified Learning (MUL) module, which unifies cross-modality labels only for modality-shared instances while retaining the original labels for modality-specific ones.
Zhiyong Li, Wei Jiang, Haojie Liu et al.· International Journal of Mac...· 0 citations
This work proposes MDCRNet, a Multi-scale Decomposed Convolution Refinement Network that enhances cross-modal feature learning and discriminative metric learning, and develops a Joint Discriminative Metric Loss incorporating a novel Granularity Discriminative Loss (GDL).
Mingsheng Zheng, Zi-Rui Jiang, Bo Liu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.