Skip to content
Open access

UniLipi: A Unified Multi-script OCR for Historical Indic Manuscripts

Aug 2026 · IEEE International Conference on Document Analysis and Recognition · pp. 140-157 · 0 citations · 65 references
Computer Science

TL;DR

UniLipi, a unified multi-script OCR model for handwritten Indic manuscripts trained jointly across 13 Indic scripts within a single framework, serves as an effective foundational pretrained model and predicts script identity and per-line native character counts, supporting practical manuscript cataloging workflows.

Abstract

Optical character recognition (OCR) for handwritten Indic manuscripts is essential for large-scale digitization and computational access to manuscript heritage. However, existing approaches are typically developed for one script at a time and require substantial script-specific customization. This limits scalability and practical deployment across diverse collections. We present UniLipi, a unified multi-script OCR model for handwritten Indic manuscripts trained jointly across 13 Indic scripts within a single framework. UniLipi directly handles realistic manuscript conditions, including extreme variation in line geometry, large variation in line length, and partial interruptions caused by non-textual manuscript entities such as holes, stains, or pictorial illustrations. To operate effectively under ultra low-resource conditions, the model leverages script-aware synthetic manuscript data generation, substantially reducing reliance on large volumes of real annotated data. Beyond historical manuscripts, we show that UniLipi serves as an effective foundational pretrained model. Specifically, its learned representations enable good OCR performance for contemporary Indic handwriting and extend to several non-Indic scripts, including Tibetan, Italian, Latin, and Chinese scripts. In addition to transcription, UniLipi predicts script identity and per-line native character counts, supporting practical manuscript cataloging workflows.

Read PDF

Similar papers

IndicPage-OCR: Robust Low-Resource Adaptation for Multi-Script Indic Handwritten Page Recognition

Experiments show substantial reductions in Word Error Rate (WER) and Character Error Rate (CER), narrowing the performance gap between commercial and freely deployable OCR systems by approximately 80% in WER and over 90% in CER, while consistently outperforming general-purpose foundation-style baselines.

Shaon Bhattacharyya, Ajoy Mondal, C. V. Jawahar et al. · 0 citations
Jul 2026

Persian Pixel: A large-scale synthetic OCR dataset for Persian language

Optical Character Recognition (OCR) for Persian remains substantially less mature than for Latin-script languages despite Persian being spoken by more than 110 million people across multiple countries. This gap arises from two fundamental challenges: the intrinsic complexity of the Perso-Arabic writing system and the l...

Pouria Mahdi, Haq Nawaz Malik · 0 citations
Open access Aug 2026

Extraction of Handwritten and Printed Cyrillic Text from Documents: A Resource-Efficient Pipeline

This paper presents a comprehensive, resource-efficient pipeline for extracting printed and handwritten Cyrillic text that integrates advanced image preprocessing, YOLO-based document structure and table recognition, and a Permuted Autoregressive Sequence (PARSeq) model trained specifically for Bulgarian.

D. Halachev, Ivan Koychev · 0 citations
Preprint Sep 2026

Rethinking Handwritten Character Recognition

Non-Latin handwritten character recognition (HCR) remains understudied. Dominant methods consider it as generic image classification, which uses model scale to implicitly learn stroke structure. Structural-prior efficiency---the principle that explicitly encoding script-geometric regularities as architectural inductive...

R. Raut, Aarav Subedi, Ashim Shrestha · 0 citations
#artificial intelligence Preprint Aug 2026

Impact of Iterative Fine-Tuning on Transcription Accuracy in Complex Historical Sanskrit Manuscripts

A local traditional OCR pipeline is introduced that can be iteratively fine-tuned on the target manuscript at the layout-level and the appearance-level, causing iterative reduction in human annotation effort, which is expensive and time-consuming as it requires historical domain expertise.

Kartik Chincholikar, Kaushik Gopalan, Mihir Hasabnis · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.