Aug 2026· International Conference on Multimedia Analysis and Pattern Recognition· pp. 724-729· 0 citations· 44 references
Abstract
Recent advances in Large Language Models (LLMs) have improved reasoning in multimodal tasks such as Visual Question Answering (VQA). However, in OCR-centric scenarios such as signboard VQA, existing approaches remain vulnerable to hallucination, weak verification, and inconsistent reasoning when integrating visual and textual cues. In this paper, we propose a graph-based multi-agent LLM framework that introduces a structured intermediate reasoning layer between perception and language reasoning. OCR entities and visual objects are organized into a lightweight, layout-aware graph, and heterogeneous LLM agents (GPT and Gemini) exchange information exclusively through this shared structure rather than through free-form text, with a dedicated verification agent scoring and pruning hypotheses against the graph before answer generation. On ViSignVQA, a Vietnamese signboard VQA benchmark, our framework attains 55.80% F1 and 23.48% EM, improving over the strongest reported baseline by 4.04 F1 and 5.40 EM points. An ablation study shows that both modalities are required: removing OCR nodes or visual nodes reduces EM to 2.09% and 2.07%, respectively. On EVJVQA, a multilingual benchmark that is not OCR-centric, the framework transfers without collapsing, reaching the second-highest F1 (0.3618) among submitted systems, although its BLEU remains low — a gap we analyze as a property of LLM-generated answers rather than of reasoning quality. Our results highlight both the value and the current limits of structured reasoning and verification in multimodal LLM systems.
ReVA is proposed, a region-aware VQA model that employs a frozen CLIP ViT-L/14 Vision Transformer (ViT) and a Qwen2.5-7B-Instruct large language model (LLM) connected through a dual bridge that aligns both whole-image and region-level representations with the LLM's embedding space.
Recent work in visual question answering has shown that vision-language models can exhibit strong reasoning capabilities by translating visual inputs into textual representations. The effectiveness of this translation depends on how well visual details are retained; models need to surface and align both explicit and im...
Ting-Chih Chen, Emile van Krieken, Shu-Jian Yu et al.· 0 citations
This work introduces DocHop, a benchmark for integrated chart--context reasoning in document-style images and constructs DocHop via a stochastic logic-first generation pipeline with controllable reasoning depth and visual density, to enable systematic evaluation.
Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park et al.· 1 citation
Large language models (LLMs) are increasingly combined with knowledge graphs (KGs) to ground reasoning in structured evidence. However, most LLM-based KGQA methods rely on traversing existing graph edges and become unreliable when reasoning paths are broken by missing facts. Alternatives that ask LLMs to generate missi...
EKS is a novel framework that leverages entity relations in commonsense knowledge graphs to dynamically generate knowledge sentences relevant to both visual and textual entities and formulates knowledge selection as a relevance scoring problem, where semantic similarity is used to measure the relevance between knowledg...
Kun Zhu, Kun Zhou, De-Xin Zhao· Multimedia Systems· 0 citations
Experiments with representative closed-source and open-source MLLMs show that OCR-grounded meta-reasoning remains far from saturated: models struggle with visible-rule application and layout-sensitive inference, while process-compliant rationales can accompany incorrect final answers under exact-match evaluation.
Geng-Xu Li, Yuan Wu, Yi Chang· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.