Skip to content
Conference

Graph-based Multi-Agent LLM Framework for OCR-based Visual Question Answering

Aug 2026 · International Conference on Multimedia Analysis and Pattern Recognition · pp. 724-729 · 0 citations · 44 references

Abstract

Recent advances in Large Language Models (LLMs) have improved reasoning in multimodal tasks such as Visual Question Answering (VQA). However, in OCR-centric scenarios such as signboard VQA, existing approaches remain vulnerable to hallucination, weak verification, and inconsistent reasoning when integrating visual and textual cues. In this paper, we propose a graph-based multi-agent LLM framework that introduces a structured intermediate reasoning layer between perception and language reasoning. OCR entities and visual objects are organized into a lightweight, layout-aware graph, and heterogeneous LLM agents (GPT and Gemini) exchange information exclusively through this shared structure rather than through free-form text, with a dedicated verification agent scoring and pruning hypotheses against the graph before answer generation. On ViSignVQA, a Vietnamese signboard VQA benchmark, our framework attains 55.80% F1 and 23.48% EM, improving over the strongest reported baseline by 4.04 F1 and 5.40 EM points. An ablation study shows that both modalities are required: removing OCR nodes or visual nodes reduces EM to 2.09% and 2.07%, respectively. On EVJVQA, a multilingual benchmark that is not OCR-centric, the framework transfers without collapsing, reaching the second-highest F1 (0.3618) among submitted systems, although its BLEU remains low — a gap we analyze as a property of LLM-generated answers rather than of reasoning quality. Our results highlight both the value and the current limits of structured reasoning and verification in multimodal LLM systems.

View source

Similar papers

#computer vision Preprint Aug 2026

ReVA: A Region-Aware Visual Assistant for Visually Grounded Question Answering

ReVA is proposed, a region-aware VQA model that employs a frozen CLIP ViT-L/14 Vision Transformer (ViT) and a Qwen2.5-7B-Instruct large language model (LLM) connected through a dual bridge that aligns both whole-image and region-level representations with the LLM's embedding space.

A. Senthil · 0 citations
#natural language process... Preprint Sep 2026

Question-Specific Knowledge Graphs for Efficient Visual Reasoning

Recent work in visual question answering has shown that vision-language models can exhibit strong reasoning capabilities by translating visual inputs into textual representations. The effectiveness of this translation depends on how well visual details are retained; models need to surface and align both explicit and im...

Ting-Chih Chen, Emile van Krieken, Shu-Jian Yu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents

This work introduces DocHop, a benchmark for integrated chart--context reasoning in document-style images and constructs DocHop via a stochastic logic-first generation pipeline with controllable reasoning depth and visual density, to enable systematic evaluation.

Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park et al. · 1 citation
#natural language process... Preprint Sep 2026

Explore-on-Graph: Hybrid Embedding-LLM Reasoning for Knowledge Graph Question Answering under Incompleteness

Large language models (LLMs) are increasingly combined with knowledge graphs (KGs) to ground reasoning in structured evidence. However, most LLM-based KGQA methods rely on traversing existing graph edges and become unreliable when reasoning paths are broken by missing facts. Alternatives that ask LLMs to generate missi...

Ola El Khatib, D. Difallah · 0 citations
Aug 2026

A visual question answering model based on entity knowledge selection

EKS is a novel framework that leverages entity relations in commonsense knowledge graphs to dynamically generate knowledge sentences relevant to both visual and textual entities and formulates knowledge selection as a relevance scoring problem, where semantic similarity is used to measure the relevance between knowledg...

Kun Zhu, Kun Zhou, De-Xin Zhao · 0 citations
#natural language process... Preprint Aug 2026

OCR-MetaReasoning Benchmark: Evaluating the Meta-Reasoning Ability of MLLMs in Text-Rich Image Understanding

Experiments with representative closed-source and open-source MLLMs show that OCR-grounded meta-reasoning remains far from saturated: models struggle with visible-rule application and layout-sensitive inference, while process-compliant rationales can accompany incorrect final answers under exact-match evaluation.

Geng-Xu Li, Yuan Wu, Yi Chang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.