Skip to content
Preprint

Thinking with Anchors: Grounded and Efficient Document Reasoning

Aug 2026 · 0 citations
Computer Science

TL;DR

ADOPD 2026 is presented, a reasoning-oriented extension of ADOPD that turns page decomposition into spatially grounded document understanding and provides a task framework that moves document understanding beyond localization toward anchor-grounded document intelligence.

Abstract

Existing document understanding benchmarks have largely focused on locating page elements, yet real-world document intelligence requires models to reason jointly about region semantics, spatial relations, and visual structure. We present ADOPD 2026, a reasoning-oriented extension of ADOPD that turns page decomposition into spatially grounded document understanding. ADOPD 2026 enriches page anchors inherited from ADOPD 2024 dataset with human-cleaned captions, semantic tags, and generated chain-of-thought (CoT) traces grounded to document regions. Instead of treating boxes, masks, and tags as independent supervision signals, we cast text blocks, visual entities, semantic labels, bounding boxes, and polygon masks as a shared vocabulary of visual anchors. This representation supports three connected capabilities. First, region-level semantic tagging asks models to identify document element types from both page context and local appearance, revealing long-tail semantic failures that standard layout benchmarks often hide. Second, unified vision-language grounding generates text regions and visual entities together with coordinates or polygonal outlines, transforming detection and segmentation outputs into structured anchors that can be reused by downstream reasoning systems. Third, current state-of-the-art models still struggle with dense counting tasks evaluated on DocCount, a benchmark derived from ADOPD 2026, highlighting the need for the Thinking-with-Anchors pipeline in document semantic understanding. By connecting page decomposition to verifiable visual-anchor reasoning, ADOPD 2026 provides a task framework that moves document understanding beyond localization toward anchor-grounded document intelligence.

View source

Similar papers

Preprint Sep 2026

Structure-Token Evidence-Anchored Reasoning for Scientific Chart Understanding

Scientific charts encode quantities in axes, legends, and geometric marks, yet large vision-language models still treat them as natural photographs. Visual in-context examples do not expose the coordinate frame; unconstrained chain-of-thought can name a plausible number that was never read from a bar. We present STEER (Structure-Token Evidence-anchored Reasoning), which freezes a Llama-3.2-Vision encoder and inserts three modules: a chart structure graph encoder (CSGE) that binds ticks, legend items, and marks; evidence-anchored step reasoning (EASR) that forces every arithmetic step to cite a graph node; and weak-parser strong-reasoner alignment (WPSR) that uses a specialized table extractor only as a teacher of node attributes. On ChartQA, STEER reaches 82.70 average relaxed accuracy versus 80.16 for ChartGemma and 76.40 for a LLaVA-CoT backbone trained on the same mix. Gains widen on CharXiv reasoning (33.60 vs. 29.20 InternVL Chat V1.5) and ChartQAPro CoT (40.70 vs. 37.17 Qwen2-VL-7B), where OCR shortcuts disappear. Ablations show that dropping node serialization or numeric candidate constraints undoes most of the reasoning lift.

Alberlucia Rafael Soarez, Camila Ferreira, Daniel Kim et al. · 0 citations
Preprint Aug 2026

FRAGMENT: Factorized Graph Representations for Document Generation and Editing via Entity-Aware Transformations

Experiments on DocLayNet, FUNSD, and SROIE evaluate FRAGMENT alongside representative autoregressive, layout-only, and graph-based baselines, providing an empirical analysis of the characteristics and trade-offs of the proposed factorized graph generation framework.

Ayoub El Bouchtili, Guilhaume Leroy-Meline · 0 citations
Book Open access Aug 2026

Docling: Converting Complex Documents into AI-Ready Structured Representations

By bridging the gap between visually complex documents and machine-readable knowledge, Docling provides a foundation for reliable document understanding in next-generation AI systems.

P. Staar · 0 citations
Preprint Aug 2026

DocTrace: Towards Traceable Long Document VQA via Hierarchical Evidence Graph Reasoning

This paper proposes DocTrace, a hierarchical framework that progressively performs evidence localization, structured document parsing, and evidence graph reasoning to enable explicit evidence provenance, and develops a two-stage training framework.

Lei Xiang, Zhi-Cheng Guan, Hong Chen et al. · 0 citations
Preprint Aug 2026

Beyond Atomic Layouts: Compositional Design Understanding with Vision-Language Models

A post-training paradigm that integrates multimodal alignment (MA) and structural perception (SP) is proposed, MA enhances element interpretation by grounding metadata-defined elements to their visual counterparts, mitigating semantic drift, and SP models layer-aware inter-element spatial relationships to improve hierarchical understanding and reduce structural ambiguity.

Yiyang Huang, Zhao-Wen Wang, Simon Jenni et al. · 0 citations

EviMap: Evidence-Grounded Hierarchical Topic Maps for Exploring Unlabeled Corpora

Research teams and organizations often explore unfamiliar free-text collections, from survey comments and reviews to reports and domain documents, before labels, queries or coding schemes exist. At this stage, the first thematic map shapes what users notice, prioritize and carry into downstream analysis, so it should be trusted only insofar as it can be verified. Existing options force a trade-off between scale and verifiability. Qualitative coding preserves evidence but is slow. Search presupposes a query. Clustering and topic models scale but produce labels users must interpret. One-shot large language model (LLM) summaries are fluent yet difficult to reproduce or audit. We present EviMap, an interactive system providing researchers and practitioners with an auditable thematic overview of such corpora. Guided by model-generated context describing the corpus and hypothesized stakeholder concerns, EviMap extracts within-document evidence phrases and organizes them, rather than whole documents, into a three-level map of aspects, groups and fine-grained topics. Embedding-based clustering narrows the search space for finer semantic judgments by the LLM. Each node traces back to supporting phrase spans, so documents link to topics through evidence they contain and users can audit labels against the original text. Users can start from a top-level corpus map, drill into topics, inspect highlighted evidence in original documents, and combine two topics to find documents discussing both. We demonstrate this workflow across six heterogeneous corpora spanning 2,108 to 101,699 documents, with a comparison against flat and hierarchical LLM baselines. By grounding every label in verbatim source spans, EviMap makes a topic map not just readable, but verifiable. Code, demo video, and interactive dashboard are available at https://github.com/zhiyintan/EviMap.

Zhiyin Tan, Changxu Duan · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.