DocIntent, a training-free Answerability-Guided Agentic Restoration framework, which first assesses question answerability, then identifies task-relevant degradations and selectively invokes restoration tools, and which consistently improves the average score and consistency of different open- and closed-source MLLMs.
Zi-Han Huang, Shi-Hang Wu, Jun-Le Liu et al.· 0 citations
TongGuOCR is proposed, a layout-aware and token-augmented multimodal large language model (MLLM) for OCR of Chinese historical documents that outperforms representative traditional task-specific OCR models, general-purpose MLLMs, and OCR-oriented MLLMs.
Zhongheng Zhou, Yi Sun, Huiguo He et al.· 0 citations
This work proposes Single-Patch Text Spotting (SPaTS), a vision-centric framework that routes each text instance through a single anchor visual token and then recovers geometry via full-image refinement and introduces Single-Patch Selective Optimization (SPaSO), a reinforcement learning framework that optimizes discrete visual-token selection using patch-level rewards.