Skip to content

MGDT: MLLM-Guided Diffusion Transformer with Relation-Adaptive Mixture-of-Experts for Multimodal Knowledge Graph Completion

Jul 2026 · arXiv.org · Vol abs/2607.15592 · 0 citations · 28 references
Computer Science

TL;DR

MGDT is proposed, a novel MKGC framework built on an align-then-diffuse paradigm that employs a Relation-Adaptive Semantic Routing Mixture-of-Experts module to select relation-relevant multimodal semantic transformation paths and suppress irrelevant modality interference.

Abstract

Multimodal Knowledge Graph Completion (MKGC) requires inferring missing entities from structural, textual, and visual cues. Existing diffusion-based MKGC methods usually denoise directly on raw multimodal features. Such a design forces the denoiser to simultaneously perform relation-dependent cue selection, cross-modal semantic alignment, and structure-aware entity generation, which introduces noisy and semantically inconsistent conditions for diffusion and consequently leads to suboptimal completion performance. To address this limitation, we propose MGDT: MLLM-Guided Diffusion Transformer with Relation-Adaptive Mixture-of-Experts (MGDT), a novel MKGC framework built on an align-then-diffuse paradigm. MGDT first employs a Relation-Adaptive Semantic Routing Mixture-of-Experts (RASR-MoE) module to select relation-relevant multimodal semantic transformation paths and suppress irrelevant modality interference. MGDT then uses a frozen Multimodal Large Language Model (MLLM) as a semantic anchor to align the routed multimodal representations into a unified latent space and reduce cross-modal semantic heterogeneity. Finally, a Knowledge Graph Diffusion Transformer (KGDT) performs graph-conditioned denoising generation in the aligned space to produce the missing entity representation. Experiments on three benchmark datasets show that MGDT consistently outperforms strong baselines.

View source

Similar papers

Open access Sep 2026

MOSAIC: A Multimodal Semantic-Oriented Alignment with Integrated Contrastive Learning for Multimodal Knowledge Graph Construction from Scientific Documents

Multimodal Knowledge Graphs (MMKGs) offer a promising paradigm for integrating heterogeneous sources into a unified, queryable, semantically structured representation. However, existing MMKG construction pipelines remain predominantly text-centric, extracting information from textual passages while leaving much of the...

Busisani Mac Dube, Jean Vincent Fonou Dombeu · 0 citations
Open access Aug 2026

HGCRec: A Heterogeneous Graph Contrastive Learning Framework for Cold-Start Multi-Modal Recommendation in Intelligent Information Systems

Personalized recommendation has become an essential component of intelligent information systems and electronic multimedia platforms. However, cold-start items with limited user–item interactions remain difficult to model, especially when collaborative signals are sparse and heterogeneous side information is underutili...

Hua-Yu Li, Yang-Chen Xu · 0 citations
2026

SIHAN: Semantic-Instance Guided Hypergraph Attention Network With Dual-View Contrastive Learning for Heterogeneous Graph

Heterogeneous graphs are well-suited to modeling the diverse types of entities and their complex interactions in the real world. However, existing Heterogeneous Graph Neural Networks (HGNNs) are typically based on the binary message-passing framework, which struggles to explicitly and finely describe the higher-order s...

Shu-Juan Wei, Hui-Jun Tang, Peng-Fei Jiao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.