Graph-based Multi-Agent LLM Framework for OCR-based Visual Question Answering
Abstract
Recent advances in Large Language Models (LLMs) have improved reasoning in multimodal tasks such as Visual Question Answering (VQA). However, in OCR-centric scenarios such as signboard VQA, existing approaches remain vulnerable to hallucination, weak verification, and inconsistent reasoning when integrating visual and textual cues. In this paper, we propose a graph-based multi-agent LLM framework that introduces a structured intermediate reasoning layer between perception and language reasoning. OCR entities and visual objects are organized into a lightweight, layout-aware graph, and heterogeneous LLM agents (GPT and Gemini) exchange information exclusively through this shared structure rather than through free-form text, with a dedicated verification agent scoring and pruning hypotheses against the graph before answer generation. On ViSignVQA, a Vietnamese signboard VQA benchmark, our framework attains 55.80% F1 and 23.48% EM, improving over the strongest reported baseline by 4.04 F1 and 5.40 EM points. An ablation study shows that both modalities are required: removing OCR nodes or visual nodes reduces EM to 2.09% and 2.07%, respectively. On EVJVQA, a multilingual benchmark that is not OCR-centric, the framework transfers without collapsing, reaching the second-highest F1 (0.3618) among submitted systems, although its BLEU remains low — a gap we analyze as a property of LLM-generated answers rather than of reasoning quality. Our results highlight both the value and the current limits of structured reasoning and verification in multimodal LLM systems.