Skip to content
Conference

Multimodal data-driven knowledge graph reasoning model for historical and cultural resources

Aug 2026 · International Conference on Advanced Manufacturing, Automation, and Deep Learning · Vol 14311, pp. 143110G - 143110G-7 · 0 citations · 10 references
Engineering

Abstract

To address the problems of heterogeneous structures, weak semantic alignment, sparse relations, and reasoning conflicts in multi-modal knowledge graph construction, this paper proposes a multi-modal data-driven knowledge graph reasoning model integrating unified encoding, graph representation learning, and rule constraints. Text descriptions, image labels, temporal attributes, and spatial coordinates are mapped into a shared representation space through modality-specific encoders and feature projection matrices. A multi-relational knowledge graph is then constructed through entity extraction, relation identification, triplet generation, and graph storage. On this basis, graph representation learning is introduced to aggregate neighborhood information and mine implicit relations among sparse nodes. To reduce semantic drift and invalid link prediction, temporal order, spatial inclusion, entity dependency, and event consistency rules are embedded into the reasoning score function. Candidate filtering and conflict resolution are further designed to improve inference efficiency under large-scale triplet conditions. Experimental results show that the proposed model achieves 0.894 Precision, 0.872 Recall, 0.883 F1, 0.907 MRR, and 0.931 Hits@10, outperforming TransE, GCN, GAT, R-GCN, and CompGCN. When the candidate triplet scale reaches 10,000, the inference response time of the proposed model is 146 ms, indicating better relation completion accuracy, reasoning stability, and inference efficiency in multi-modal knowledge graph reasoning tasks.

View source