Aug 2026· IEEE International Conference on Document Analysis and Recognition· pp. 140-157· 0 citations· 65 references
Computer Science
TL;DR
UniLipi, a unified multi-script OCR model for handwritten Indic manuscripts trained jointly across 13 Indic scripts within a single framework, serves as an effective foundational pretrained model and predicts script identity and per-line native character counts, supporting practical manuscript cataloging workflows.
Abstract
Optical character recognition (OCR) for handwritten Indic manuscripts is essential for large-scale digitization and computational access to manuscript heritage. However, existing approaches are typically developed for one script at a time and require substantial script-specific customization. This limits scalability and practical deployment across diverse collections. We present UniLipi, a unified multi-script OCR model for handwritten Indic manuscripts trained jointly across 13 Indic scripts within a single framework. UniLipi directly handles realistic manuscript conditions, including extreme variation in line geometry, large variation in line length, and partial interruptions caused by non-textual manuscript entities such as holes, stains, or pictorial illustrations. To operate effectively under ultra low-resource conditions, the model leverages script-aware synthetic manuscript data generation, substantially reducing reliance on large volumes of real annotated data. Beyond historical manuscripts, we show that UniLipi serves as an effective foundational pretrained model. Specifically, its learned representations enable good OCR performance for contemporary Indic handwriting and extend to several non-Indic scripts, including Tibetan, Italian, Latin, and Chinese scripts. In addition to transcription, UniLipi predicts script identity and per-line native character counts, supporting practical manuscript cataloging workflows.
Experiments show substantial reductions in Word Error Rate (WER) and Character Error Rate (CER), narrowing the performance gap between commercial and freely deployable OCR systems by approximately 80% in WER and over 90% in CER, while consistently outperforming general-purpose foundation-style baselines.
Shaon Bhattacharyya, Ajoy Mondal, C. V. Jawahar et al.· 0 citations
Optical Character Recognition (OCR) for Persian remains substantially less mature than for Latin-script languages despite Persian being spoken by more than 110 million people across multiple countries. This gap arises from two fundamental challenges: the intrinsic complexity of the Perso-Arabic writing system and the l...
Pouria Mahdi, Haq Nawaz Malik· arXiv.org· 0 citations
This paper presents a comprehensive, resource-efficient pipeline for extracting printed and handwritten Cyrillic text that integrates advanced image preprocessing, YOLO-based document structure and table recognition, and a Permuted Autoregressive Sequence (PARSeq) model trained specifically for Bulgarian.
D. Halachev, Ivan Koychev· Computational Linguistics in...· 0 citations
OmniHandwritingOCR provides a challenging testbed for diagnosing language, content, structural, and visual-grounding failure modes of multimodal models in handwritten OCR scenarios.
Non-Latin handwritten character recognition (HCR) remains understudied. Dominant methods consider it as generic image classification, which uses model scale to implicitly learn stroke structure. Structural-prior efficiency---the principle that explicitly encoding script-geometric regularities as architectural inductive...
R. Raut, Aarav Subedi, Ashim Shrestha· 0 citations
A local traditional OCR pipeline is introduced that can be iteratively fine-tuned on the target manuscript at the layout-level and the appearance-level, causing iterative reduction in human annotation effort, which is expensive and time-consuming as it requires historical domain expertise.