Skip to content

ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains

Sep 2026 · Lecture Notes in Computer Science · pp. 336-352 · 1 citation · 14 references
Computer Science

TL;DR

The results show that the strongest systems move beyond single-pass prompting and instead rely on structured evidence extraction, retrieval, verification, and orchestration across multiple components.

Abstract

In this report we present results of the ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains. This competition aimed to advance research in document understanding through the task of Visual Question Answering (VQA). Building upon previous DocVQA benchmarks, this competition introduces challenging reasoning questions over a diverse collection of documents spanning eight domains, including business reports, scientific papers, slides, posters, maps, comics, infographics, and engineering drawings. The competition concluded with 20 valid submissions from 8 teams spanning zero-shot VLMs, OCR and parser-augmented pipelines, agentic retrieval systems, multi-agent ensembles, and fine-tuned multimodal models. The results show that the strongest systems move beyond single-pass prompting and instead rely on structured evidence extraction, retrieval, verification, and orchestration across multiple components.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents

This work introduces DocHop, a benchmark for integrated chart--context reasoning in document-style images and constructs DocHop via a stochastic logic-first generation pipeline with controllable reasoning depth and visual density, to enable systematic evaluation.

Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park et al. · 1 citation

RAG Pipeline Strategies for Ukrainian Multi-Domain Document Understanding Task

This work developed a modular retrieval-augmented generation (RAG) pipeline and conducted a series of ablation experiments over its individual components to identify the best-performing strategy at each stage, demonstrating that isolated curation of RAG components can yield strong performance for Ukrainian document gro...

Mykola Nosenko, Pavlo Kilko, Markus Reuter et al. · 0 citations
#computer vision Preprint Sep 2026

Evolution of Multimodal Question Answering: From Modality-Adaptive Extraction to Unified Language Representation

The rapid growth of multimodal data has intensified the need for question answering (QA) systems capable of reasoning across heterogeneous sources such as text, tables, and images. In this paper, we present a comprehensive methodological comparison of three influential frameworks, namely Multimodal Adaptive Extraction...

Abdullah Al Shafi · 0 citations
#artificial intelligence Preprint Aug 2026

CrossModalQA: A Cross-modal and Multi-hop Benchmark for Multimodal Retrieval-augmented Generation

CrossModalQA is introduced, an open-domain benchmark for evaluating multimodal retrieval and reasoning over heterogeneous corpora and it is revealed that complete cross-modal retrieval contributes more to answer accuracy than generator scaling, while multi-image retrieval and reasoning remain the primary bottlenecks li...

Jia-Cheng Cai, Zi-Jin Hong, Zheng Yuan et al. · 0 citations
Preprint Aug 2026

BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models

BEAR-Bench (Bilingual Enterprise and Academic Reasoning), a self-contained, complex English-and-Russian benchmark comprising 1000 human-annotated questions based on text-rich business and scientific documents, is introduced, and existing hallucination detection methods are compared.

L. Chubarova, A. Kuleshova, D. P. Volkov et al. · 0 citations
Conference Aug 2026

Graph-based Multi-Agent LLM Framework for OCR-based Visual Question Answering

Recent advances in Large Language Models (LLMs) have improved reasoning in multimodal tasks such as Visual Question Answering (VQA). However, in OCR-centric scenarios such as signboard VQA, existing approaches remain vulnerable to hallucination, weak verification, and inconsistent reasoning when integrating visual and...

T. Dang, H. Nguyen, Kiet Van Nguyen · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.