Aug 2026· Electronics· Vol 15, pp. 3414· 0 citations· 29 references
TL;DR
This work proposes Collaborative Entity Linking through Multichannel Interaction (CELMI), a four-channel framework that jointly models textual semantics, visual perception, cross-modal alignment, and inter-mention collaboration and introduces a multi-channel consistency objective that combines per-channel contrastive losses with an overall ranking loss.
Abstract
Multimodal entity linking (MEL) grounds entity mentions in text-image contexts to entries in a structured knowledge base; however, most existing systems still decompose a multimodal document into independent mention-level decisions. This formulation overlooks a central tension of real-world MEL: the evidence needed to disambiguate an ambiguous mention is often distributed across co-occurring mentions, visual context, and cross-modal consistency, rather than being contained in the mention itself. To address this limitation, we propose Collaborative Entity Linking through Multichannel Interaction (CELMI), a four-channel framework that jointly models textual semantics, visual perception, cross-modal alignment, and inter-mention collaboration. CELMI employs dual-level textual alignment, text-guided visual gating, learned cross-modal projection, and attention-based entity graph propagation. To stabilize joint optimization, we further introduce a multi-channel consistency objective that combines per-channel contrastive losses with an overall ranking loss, reducing channel dominance and representation collapse. Among conventional non-LLM/VLM MEL models, CELMI achieves the strongest MRR and Hits@1 performance on WikiMEL, RichpediaMEL, and WikiDiverse, with Hits@1 scores of 89.03% on WikiMEL and 83.02% on RichpediaMEL; LLM/VLM systems remain stronger on WikiDiverse, positioning CELMI as a lightweight complement. Progressive stress ablation shows a 38.50 percentage-point absolute Hits@1 drop when the architecture is reduced to a single unimodal endpoint, confirming that the gains arise from synergistic channel interaction rather than isolated module effects.
This work introduces MMHRAG, a novel Multi-Modal Hierarchical Retrieval-Augmented Generation framework that achieves cross-modal interaction on DocQA for the first time, and designs a Summarizing Agent to resolve logical conflicts, information redundancy, and granularity discrepancies among retrieved cross-modal eviden...
Jia-Yuan Wang, Jie Lian, Fu Zhao et al.· Proceedings of the 32nd ACM...· 0 citations
CrossModalQA is introduced, an open-domain benchmark for evaluating multimodal retrieval and reasoning over heterogeneous corpora and it is revealed that complete cross-modal retrieval contributes more to answer accuracy than generator scaling, while multi-image retrieval and reasoning remain the primary bottlenecks li...
Jia-Cheng Cai, Zi-Jin Hong, Zheng Yuan et al.· 0 citations
PrismF is a unified framework that synergizes multi-perspective enhancement with progressive fusion to extract stronger signals from diverse inputs and improves cross-modal integration through a progressive fusion strategy that dynamically calibrates inter-modal interactions.
Chen-Yi Xiong, Yan Zhang, Jing Hu et al.· 0 citations
Visual Question Answering (VQA) aims to achieve cross-modal semantic understanding through joint modeling of visual content and natural language. Although existing attention-based approaches effectively align features, they struggle to simultaneously accommodate global semantic modeling and local fine-grained perceptio...
Gan-Long Zhou, Dezhi Han, Xiang Shen et al.· Computer Science and Informa...· 0 citations
A novel framework for constructing a Context-Enhanced MMKG (CEMMKG) is proposed, effective in leveraging contextual information to improve MMKG-based RAG performance and its effectiveness across different MMKG-based RAG methods demonstrates its broad applicability.
Zongyu Wu, Yilong Wang, Xiaochen Wang et al.· 0 citations
Multimodal Knowledge Graphs (MMKGs) offer a promising paradigm for integrating heterogeneous sources into a unified, queryable, semantically structured representation. However, existing MMKG construction pipelines remain predominantly text-centric, extracting information from textual passages while leaving much of the...
Busisani Mac Dube, Jean Vincent Fonou Dombeu· Big Data and Cognitive Compu...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.