Skip to content
Book Open access

Diagnosing Evidence Utilization in Multimodal Document Question Answering

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · 0 citations · 35 references

Abstract

Recent Multimodal Large Language Models (MLLMs) support retrieval-augmented generation (RAG) for document question answering (QA), yet it remains unclear how effectively they use the provided evidence during answer generation. In this work, we conduct a controlled empirical study of 7 popular MLLMs on long multimodal multi-document question answering in a RAG setting. Our analysis shows that zero-shot performance varies substantially across evidence types (e.g., text, image, table, chart, and cross-evidence), revealing a strong reliance on text-based evidence and weaker performance on image-only and cross-evidence inputs. We further find that supervised finetuning yields limited and dataset-dependent improvements, often preserving existing evidence-type disparities. To better understand these behaviours, we perform attention-based analysis to examine how models allocate attention across different token types (image evidence, text evidence, system prompt, question, and output tokens), and find that low attention allocation to image tokens is associated with weaker performance on image-based evidence. Our findings provide insights into how effectively current MLLMs use multimodal evidence in document QA and highlight the key limitations in multimodal document understanding.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.