Skip to content

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Book Open access Aug 2026

End-to-End Structured Information Extraction from Mixed-Script Documents in Open Compositional Symbol Systems

Structurally extracting information from mixed-script documents that interleave standard text with open, compositional symbol systems is challenging for both optical character recognition (OCR) and vision–language models (VLMs). This difficulty is epitomized by Jianzi Pu—the ancient Guqin tablature. Unlike closed-set scripts, Jianzi glyphs are formed via infinite compositional rules and are densely integrated with Hanzi text, demanding a model that can simultaneously perform script discrimination, layout parsing, and structural transcription. We propose JZ-Tab, the first framework dedicated to the automated recognition of Jianzi Pu, which functions as an end-to-end structured visual information extraction system for mixed-script documents. Unlike traditional pipelines, JZ-Tab generates layout-aware markup directly from full-page images, bypassing the need for pre-segmentation. Specifically, to overcome the total absence of large-scale annotated datasets, we develop a novel, scalable synthetic-to-real pipeline that constructs layout-consistent pages from canonicalized glyph inventories. Furthermore, to capture the unique action-oriented semantics of the tablature, we introduce music-structured generation, injecting sequential regularities derived from symbolic music logic into the learning process. Finally, we train a VLM for direct page-to-markup generation. Evaluated zero-shot on authentic historical manuscripts, JZ-Tab improves F1 by +40.10 over the strongest generic VLM baselines, highlighting its potential for large-scale automated digitization of historical Guqin manuscripts and open, compositional symbol systems.

Zehan Li, Fu Zhang, Zhijun Liu et al. · 0 citations