Open access
Jul 2026
Construction of a fine-grained retrieval model for archival text-image based on the integration of scene graph generation and attention mechanism
A five-tier architectural model is devised that incorporates a dedicated scene graph generation module tailored for archival data, aiming to enhance element detection and three-tier attention fusion module that integrates scene graph, text, and cross-modal features to ensure precise feature alignment.
Mengyuan Zhang
· PLoS ONE · 0 citations