Collaborative Multimodal Entity Linking via Multi-Channel Neural Cross-Modal Interaction with Synergistic Consistency Optimization
This work proposes Collaborative Entity Linking through Multichannel Interaction (CELMI), a four-channel framework that jointly models textual semantics, visual perception, cross-modal alignment, and inter-mention collaboration and introduces a multi-channel consistency objective that combines per-channel contrastive l...