Jul 2026· IEEE transactions on circuits and systems for video technology (Print)· Vol abs/2607.29207, pp. 1-1· 0 citations· 72 references
Computer Science
TL;DR
A robust multi-modal ReID framework with dual semantic guidance and global-local mutual modulation, which mainly consists of three key components, namely the Text-Semantic Injector (TSI), the Masked Global-Local Modulator (MGLM), and the Hierarchical MoE Fusion (HMF).
Abstract
Multi-modal object Re-Identification (ReID) aims to retrieve target instances by leveraging complementary information across modalities. However, existing methods suffer from two challenges. First, they often fail to exploit well-aligned and reliable semantic priors, making them vulnerable to background clutter and cross-modal misalignment. On the other hand, they typically rely on holistic feature modeling, overlooking the synergy between global and local representations. To overcome these limitations, we propose a robust multi-modal ReID framework with dual semantic guidance and global-local mutual modulation, which mainly consists of three key components, namely the Text-Semantic Injector (TSI), the Masked Global-Local Modulator (MGLM), and the Hierarchical MoE Fusion (HMF). The TSI enhances semantic awareness by integrating clean and coherent textual features into visual tokens. The MGLM enables part-aware cross-modal interaction through joint guidance from soft masks and global context, improving fine-grained feature alignment. Finally, the HMF adaptively aggregates multi-spectral features under local semantic supervision, yielding discriminative and robust representations. Extensive experiments on three multi-modal ReID benchmarks demonstrate the effectiveness of the proposed method. The code will be made publicly available at https://github.com/zw-absin/DSGM upon acceptance.
Multi-modal object Re-Identification (ReID) benefits from complementary information across heterogeneous imaging modalities. To further enrich semantic representation, text descriptions have recently been incorporated as an additional modality. However, recent vision-language approaches often treat text descriptions as...
Wei-Xiang Zhou, Yu-Hao Wang, Xing-Guo Xu et al.· 0 citations
Multi-modal object re-identification (Re-ID) aims to facilitate cross-camera object retrieval in complex environments by leveraging complementary information from visual (e.g., RGB, NIR, TIR) and textual modalities. However, existing approaches often lack principled feature disentanglement and coherent multi-modal inte...
Cheng Huang, Jun-Jie Huang, Long Lan et al.· 0 citations
Single-modal object Re-identification (Re-ID) frequently suffers from performance degradation under complex and dynamic scene conditions. Although multi-modal object Re-ID leverages complementary information across diverse modalities to alleviate this issue, existing approaches remain highly susceptible to irrelevant b...
Lin Qi, Tian-Cun Guo, Yan-Fei Dong et al.· Electronics· 0 citations
Visible-infrared person re-identification (VI-ReID) is an important technique for around-the-clock person matching, and its primary challenge arises from substantial cross-modal discrepancies. To address this challenge, we propose a Decoupled Information-Guided Cross-Modal Alignment (DIGCA) framework organized into thr...
The visual modality, i.e., images, plays a key role in multi-modal entity alignment (MMEA). Existing approaches often directly fuse the image with other modalities to align different entities. Although simple, such strategies overlook the potential noise in the images and their semantic misalignment with corresponding...
Chen-Xiao Li, Yun-He Feng, Dongfang Liu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.