Preprint
Aug 2026
Query-Driven Multimodal Information Extraction from Long Documents
This work proposes query-driven image-text joint extraction from long documents, requiring models to output query-requested textual attribute values and corresponding image bounding boxes, and designed a two-level taxonomy that operates at the query and instance levels.
Yi-Zhou Gao, Ding Xia, Xi Yang
· 0 citations