Skip to content

Author

Peirong Zhang

We have 3 of 25 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Aug 2026

DocIntent: Answerability-Guided Agentic Restoration for Real-World Document Visual Question Answering

DocIntent, a training-free Answerability-Guided Agentic Restoration framework, which first assesses question answerability, then identifies task-relevant degradations and selectively invokes restoration tools, and which consistently improves the average score and consistency of different open- and closed-source MLLMs.

Zi-Han Huang, Shi-Hang Wu, Jun-Le Liu et al. · 0 citations
Preprint Aug 2026

TongGuOCR: A Layout-Aware and Token-Augmented OCR MLLM for Chinese Historical Documents

TongGuOCR is proposed, a layout-aware and token-augmented multimodal large language model (MLLM) for OCR of Chinese historical documents that outperforms representative traditional task-specific OCR models, general-purpose MLLMs, and OCR-oriented MLLMs.

Zhongheng Zhou, Yi Sun, Huiguo He et al. · 0 citations
Jul 2026

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting

This work proposes Single-Patch Text Spotting (SPaTS), a vision-centric framework that routes each text instance through a single anchor visual token and then recovers geometry via full-image refinement and introduces Single-Patch Selective Optimization (SPaSO), a reinforcement learning framework that optimizes discrete visual-token selection using patch-level rewards.

Rui Tang, Wentao Yang, Peirong Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.