Skip to content

Infinity-Parser2 Technical Report

Jul 2026 · arXiv.org · Vol abs/2607.07836 · 1 citation · 84 references
Computer Science

TL;DR

This work presents Infinity-Parser2, a large multimodal model that couples a controllable data-synthesis pipeline with multi-task reinforcement learning for end-to-end document parsing, addressing the persistent scarcity of faithfully annotated parsing corpora.

Abstract

We present Infinity-Parser2, a large multimodal model that couples a controllable data-synthesis pipeline with multi-task reinforcement learning for end-to-end document parsing, addressing the persistent scarcity of faithfully annotated parsing corpora. Our contributions are threefold. First, we build a scalable synthesis engine, pairing a controllable rendering framework with an iterative refinement loop, and use it to construct and open-source Infinity-Doc2-5M: a 5-million-sample bilingual (Chinese/English) corpus spanning diverse document types, annotated with element bounding boxes, canonical content forms (Markdown, HTML, LaTeX, SMILES, structured charts), and full-page reading order. Second, we introduce a verifiable, multi-task reward system that enables Joint Reinforcement Learning across eight co-trained objectives (document parsing, layout analysis, table parsing, math formula parsing, chart parsing, chemical formula parsing, document VQA, and general multimodal understanding), unifying perception, structure, and reasoning in a single optimization signal. Third, we release two variants under a shared architecture: Infinity-Parser2-Flash, optimized for low-latency inference with a 3.68x throughput gain over Infinity-Parser-7B, and Infinity-Parser2-Pro, engineered for precision-critical settings. Infinity-Parser2-Pro reaches state-of-the-art 87.6% on olmOCR-Bench and 74.3% on ParseBench, surpassing DeepSeek-OCR-2, PaddleOCR-VL-1.5, and MinerU2.5, with strong generalization to charts, chemical formulas, and document VQA.

View source

Similar papers

Review Aug 2026

FinixDoc: Rethinking Financial Document Parsing Beyond Saturated Benchmarks

This report presents FinixDoc, an end-to-end agentic parsing system for real-world financial documents, with FinixDoc-VL, a 4B-scale vision-language model built on Qwen3-VL-4B, as its core parser, and introduces a Document Parsing Capability Matrix organized along two practical axes: visual quality and document scale.

Hang Wang, Jin Zhang, Guoliang Xu et al. · 1 citation
Book Open access Aug 2026

UniDocVLM: Enhancing Visual Reasoning and Document Understanding for VLM via Reinforcement Learning

Document question answering over scanned pages requires two coupled abilities: (i) canonicalizing complex layouts into a faithful textual structure, and (ii) selecting and reasoning over query-relevant evidence from that structure. Most existing pipelines decouple OCR from retrieval-augmented reasoning and optimize OCR...

Zong-Sheng Cao, Anran Liu, Jun Xie et al. · 0 citations
#artificial intelligence Preprint Aug 2026

NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference

Multimodal models often build on architectures designed for generative vision-language modeling, typically combining separately pretrained vision encoders with causal language models. Visual document retrievers such as ColPali repurpose these models as encoders, carrying over the parameter and compute overhead of a VLM...

Aur'elien Lac, Tony Wu · 0 citations
#machine learning Preprint Aug 2026

Can LLMs Use Relational Transformer Embeddings?

It is argued that soft-token fusion requires stronger alignment objectives and schema-aware design before it can serve as a reliable route to relational prediction.

Francisco Galuppo Azevedo, Clarissa Lima Loures · 0 citations
Jul 2026

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning

Trace is introduced, a taxonomy-guided environment for multidomain visual reasoning that factorizes task construction into a scene grammar and an executable task program, separating visual realization from answer computation, providing evidence that broad procedural training can transfer beyond the generated task distr...

Md Tanvirul Alam · 0 citations
Jul 2026

LatentMT: Machine Translation with Latent Reasoning

LatentMT is introduced, the first systematic study of latent-reasoning LoopLMs for machine translation that adapts a small 2.6B-parameter backbone model with lightweight training and shows that hidden-representation differences shrink along the recurrent reasoning-step axis, supporting the observed saturation in perfor...

Wei-Rui Chen, Samar M. Magdy, Chiyu Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.