This paper presents a novel approach termed Intermediate Shared Feature Network (ISFNet) that explicitly addresses cross-modal person re-identification between visible and infrared domains by exploiting intermediate feature representations within a dual-stream backbone.
Abstract
Cross-modal person re-identification between visible and infrared domains remains a challenging problem due to significant modality gaps. This paper presents a novel approach termed Intermediate Shared Feature Network (ISFNet) that explicitly addresses this issue by exploiting intermediate feature representations within a dual-stream backbone. Unlike conventional methods that primarily focus on final-layer features, ISFNet introduces two complementary components: a Multi-layer Feature Cascade Module (MFCM) that aggregates discriminative features across different network stages, and a Dual Feature Generation Module (DFGM) that creates diverse intermediate representations through Instance-Batch Normalization variants. By integrating these modules, ISFNet effectively bridges the cross-modal gap and improves matching accuracy. Comprehensive experiments on the SYSU-MM01 and RegDB datasets demonstrate that the proposed method achieves competitive performance against state-of-the-art approaches, with noticeable improvements in both Rank-1 accuracy and mean average precision.
This work proposes MDCRNet, a Multi-scale Decomposed Convolution Refinement Network that enhances cross-modal feature learning and discriminative metric learning, and develops a Joint Discriminative Metric Loss incorporating a novel Granularity Discriminative Loss (GDL).
Mingsheng Zheng, Zi-Rui Jiang, Bo Liu et al.· 0 citations
This study proposes a novel dual-path convolution based multi-scale feature alignment (DCMFA) network that significantly outperforms existing mainstream methods in terms of recognition accuracy.
Bailiang Huang, Bin Chen, Tian-Ran Sun et al.· 電腦學刊· 0 citations
It is argued that VI-ReID should be treated as an early cross-modal correspondence learning problem rather than only a late embedding alignment problem, and CMIA-Net is proposed, a framework that establishes bidirectional visible-infrared interaction at shallow backbone stages and introduces Spectral-Invariant Augmenta...
Dao-Li Zhang, Qi-Cheng Liu· Engineering Research Express· 0 citations
A Progressively Biased Split Vision Transformer (PBSVT) is proposed, which combines a split ViT backbone with progressive bias training to gradually reduce RGB-dominant bias while preserving modality-shared structure and demonstrates the effectiveness of progressive modality transition for robust VI-ReID representation...
Mengru Jiao, Xin-Yue Xu, Jun-Feng Zhang· International journal of pat...· 0 citations
Single-modal object Re-identification (Re-ID) frequently suffers from performance degradation under complex and dynamic scene conditions. Although multi-modal object Re-ID leverages complementary information across diverse modalities to alleviate this issue, existing approaches remain highly susceptible to irrelevant b...
Lin Qi, Tian-Cun Guo, Yan-Fei Dong et al.· Electronics· 0 citations
As a critical task in intelligent surveillance and smart city systems, person re-identification (ReID) addresses the challenge of matching individuals across non-overlapping camera views. Although fully unsupervised learning (USL) methods based on pseudo-label training have achieved remarkable progress, they still face...
Qing Tian, Bing-Hui Zhang, Bin Wang et al.· IEEE Transactions on Informa...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.