Skip to content

Author

Yikun Wang

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents

Document parsing aims to transform unstructured documents into structured and machine-readable representations. Recent advances in Vision-Language Models (VLMs) have significantly advanced document parsing. However, existing approaches still face two major challenges. First, decoupled VLM-based methods heavily rely on accurate layout analysis, where geometric distortions in camera-captured documents can introduce cascading errors. Second, although end-to-end VLM-based methods alleviate the dependence on explicit layout detection, they often suffer from redundant generation, hallucinations, and insufficient structural reasoning in high-resolution scenarios. To address these challenges, we propose NaviDC-OCR, a unified framework for document parsing. NaviDC-OCR introduces deformation-aware learning to incorporate geometric perception into VLMs and proposes an adaptive sampling mechanism for complex layout representation. Furthermore, a content-structure decoupled learning strategy is developed to explicitly model formula grammars and table structures, enabling more effective structured representation learning. Extensive experiments demonstrate that NaviDC-OCR achieves state-of-the-art performance across diverse document parsing benchmarks. It obtains overall scores of 96.87, 88.53 and 78.41 on OmniDocBench v1.6, Wild-OmniDocBench, and PureDocBench, respectively, and ranks first in the ICDAR 2026 Sci-ImageMiner Challenge. These results validate the effectiveness and generalization capability of NaviDC-OCR in complex document parsing scenarios.

Peng Cai, Zhaofan Zou, Shifa Liu et al. · 0 citations
Preprint Jul 2026

DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data

DataOrchestra, a framework that unifies different processing operations and orchestrates an example-specific pipeline for each example, is proposed and effective for math continued pretraining and outperforms stronger processing baselines, while reducing processing compute by skipping unnecessary downstream operations.

Zhen Huang, Yikun Wang, Shijie Xia et al. · 0 citations
Jun 2026

Diagnosing and Mitigating Context Rot in Long-horizon Search

Through a systematic study of four flagship models, a previously overlooked phenomenon is identified: under extensive context, models give up or provide uncertain incorrect answers long before exhausting the context window.

Shijie Xia, Yikun Wang, Zhen Huang et al. · 2 citations