Skip to content
Open access

Collaborative Multimodal Entity Linking via Multi-Channel Neural Cross-Modal Interaction with Synergistic Consistency Optimization

Aug 2026 · Electronics · Vol 15, pp. 3414 · 0 citations · 29 references

TL;DR

This work proposes Collaborative Entity Linking through Multichannel Interaction (CELMI), a four-channel framework that jointly models textual semantics, visual perception, cross-modal alignment, and inter-mention collaboration and introduces a multi-channel consistency objective that combines per-channel contrastive losses with an overall ranking loss.

Abstract

Multimodal entity linking (MEL) grounds entity mentions in text-image contexts to entries in a structured knowledge base; however, most existing systems still decompose a multimodal document into independent mention-level decisions. This formulation overlooks a central tension of real-world MEL: the evidence needed to disambiguate an ambiguous mention is often distributed across co-occurring mentions, visual context, and cross-modal consistency, rather than being contained in the mention itself. To address this limitation, we propose Collaborative Entity Linking through Multichannel Interaction (CELMI), a four-channel framework that jointly models textual semantics, visual perception, cross-modal alignment, and inter-mention collaboration. CELMI employs dual-level textual alignment, text-guided visual gating, learned cross-modal projection, and attention-based entity graph propagation. To stabilize joint optimization, we further introduce a multi-channel consistency objective that combines per-channel contrastive losses with an overall ranking loss, reducing channel dominance and representation collapse. Among conventional non-LLM/VLM MEL models, CELMI achieves the strongest MRR and Hits@1 performance on WikiMEL, RichpediaMEL, and WikiDiverse, with Hits@1 scores of 89.03% on WikiMEL and 83.02% on RichpediaMEL; LLM/VLM systems remain stronger on WikiDiverse, positioning CELMI as a lightweight complement. Progressive stress ablation shows a 38.50 percentage-point absolute Hits@1 drop when the architecture is reduced to a single unimodal endpoint, confirming that the gains arise from synergistic channel interaction rather than isolated module effects.

Read PDF

Similar papers

Book Open access Aug 2026

Multi-Modal Hierarchical Retrieval-Augmented Generation for Document Question Answering

This work introduces MMHRAG, a novel Multi-Modal Hierarchical Retrieval-Augmented Generation framework that achieves cross-modal interaction on DocQA for the first time, and designs a Summarizing Agent to resolve logical conflicts, information redundancy, and granularity discrepancies among retrieved cross-modal eviden...

Jia-Yuan Wang, Jie Lian, Fu Zhao et al. · 0 citations
#artificial intelligence Preprint Aug 2026

CrossModalQA: A Cross-modal and Multi-hop Benchmark for Multimodal Retrieval-augmented Generation

CrossModalQA is introduced, an open-domain benchmark for evaluating multimodal retrieval and reasoning over heterogeneous corpora and it is revealed that complete cross-modal retrieval contributes more to answer accuracy than generator scaling, while multi-image retrieval and reasoning remain the primary bottlenecks li...

Jia-Cheng Cai, Zi-Jin Hong, Zheng Yuan et al. · 0 citations
#artificial intelligence Preprint Aug 2026

More Perspectives, Stronger Signals: Multi-Perspective Enhancement and Progressive Fusion for Multimodal Entity Representation Learning

PrismF is a unified framework that synergizes multi-perspective enhancement with progressive fusion to extract stronger signals from diverse inputs and improves cross-modal integration through a progressive fusion strategy that dynamically calibrates inter-modal interactions.

Chen-Yi Xiong, Yan Zhang, Jing Hu et al. · 0 citations
Open access 2026

Decoupled global-local collaborative network for visual question answering

Visual Question Answering (VQA) aims to achieve cross-modal semantic understanding through joint modeling of visual content and natural language. Although existing attention-based approaches effectively align features, they struggle to simultaneously accommodate global semantic modeling and local fine-grained perceptio...

Gan-Long Zhou, Dezhi Han, Xiang Shen et al. · 0 citations
Preprint Aug 2026

Multi-Granularity Context-Enhanced RAG over Multimodal Knowledge Graphs

A novel framework for constructing a Context-Enhanced MMKG (CEMMKG) is proposed, effective in leveraging contextual information to improve MMKG-based RAG performance and its effectiveness across different MMKG-based RAG methods demonstrates its broad applicability.

Zongyu Wu, Yilong Wang, Xiaochen Wang et al. · 0 citations
Open access Sep 2026

MOSAIC: A Multimodal Semantic-Oriented Alignment with Integrated Contrastive Learning for Multimodal Knowledge Graph Construction from Scientific Documents

Multimodal Knowledge Graphs (MMKGs) offer a promising paradigm for integrating heterogeneous sources into a unified, queryable, semantically structured representation. However, existing MMKG construction pipelines remain predominantly text-centric, extracting information from textual passages while leaving much of the...

Busisani Mac Dube, Jean Vincent Fonou Dombeu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.