Sep 2026· Lecture Notes in Computer Science· pp. 336-352· 1 citation· 14 references
Computer Science
TL;DR
The results show that the strongest systems move beyond single-pass prompting and instead rely on structured evidence extraction, retrieval, verification, and orchestration across multiple components.
Abstract
In this report we present results of the ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains. This competition aimed to advance research in document understanding through the task of Visual Question Answering (VQA). Building upon previous DocVQA benchmarks, this competition introduces challenging reasoning questions over a diverse collection of documents spanning eight domains, including business reports, scientific papers, slides, posters, maps, comics, infographics, and engineering drawings. The competition concluded with 20 valid submissions from 8 teams spanning zero-shot VLMs, OCR and parser-augmented pipelines, agentic retrieval systems, multi-agent ensembles, and fine-tuned multimodal models. The results show that the strongest systems move beyond single-pass prompting and instead rely on structured evidence extraction, retrieval, verification, and orchestration across multiple components.
This work introduces DocHop, a benchmark for integrated chart--context reasoning in document-style images and constructs DocHop via a stochastic logic-first generation pipeline with controllable reasoning depth and visual density, to enable systematic evaluation.
Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park et al.· 1 citation
This work developed a modular retrieval-augmented generation (RAG) pipeline and conducted a series of ablation experiments over its individual components to identify the best-performing strategy at each stage, demonstrating that isolated curation of RAG components can yield strong performance for Ukrainian document gro...
Mykola Nosenko, Pavlo Kilko, Markus Reuter et al.· 0 citations
The rapid growth of multimodal data has intensified the need for question answering (QA) systems capable of reasoning across heterogeneous sources such as text, tables, and images. In this paper, we present a comprehensive methodological comparison of three influential frameworks, namely Multimodal Adaptive Extraction...
CrossModalQA is introduced, an open-domain benchmark for evaluating multimodal retrieval and reasoning over heterogeneous corpora and it is revealed that complete cross-modal retrieval contributes more to answer accuracy than generator scaling, while multi-image retrieval and reasoning remain the primary bottlenecks li...
Jia-Cheng Cai, Zi-Jin Hong, Zheng Yuan et al.· 0 citations
BEAR-Bench (Bilingual Enterprise and Academic Reasoning), a self-contained, complex English-and-Russian benchmark comprising 1000 human-annotated questions based on text-rich business and scientific documents, is introduced, and existing hallucination detection methods are compared.
L. Chubarova, A. Kuleshova, D. P. Volkov et al.· 0 citations
Recent advances in Large Language Models (LLMs) have improved reasoning in multimodal tasks such as Visual Question Answering (VQA). However, in OCR-centric scenarios such as signboard VQA, existing approaches remain vulnerable to hallucination, weak verification, and inconsistent reasoning when integrating visual and...
T. Dang, H. Nguyen, Kiet Van Nguyen· International Conference on...· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.