Skip to content
Review Open access

Navigating the Digitization Gap: An Indirect Evidence Synthesis of AI Methods for Low-Resource Chagatai Manuscripts

Jul 2026 · Information · Vol 17, pp. 681 · 0 citations · 68 references

TL;DR

A systematic literature review following the PRISMA guidelines to examine artificial intelligence methods for handwritten text recognition (HTR) and text restoration in low-resource languages and proposes a concrete development roadmap focusing on systematic digitization, expert annotation, transfer learning, and the creation of baseline models to enable reproducible evaluations.

Abstract

Many historical handwritten records in low-resource languages remain difficult to access through modern digital systems. This limits efforts to preserve and study cultural heritage at scale. Chagatai manuscripts exemplify these challenges within the Eastern Turki tradition. For centuries, it served as a major written language across Central Asia and supported a rich literary tradition. Large collections of Chagatai manuscripts still survive today, yet only a small amount of this material exists in digital form. As the technical literature specifically focused on Chagatai-HTR remains in its nascent stage, this review synthesizes indirect evidence from taxonomically related Perso-Arabic scripts to establish a foundational research framework. This article presents a systematic literature review following the PRISMA guidelines to examine artificial intelligence methods for handwritten text recognition (HTR) and text restoration in low-resource languages. Analyzing 50 studies published between 2020 and 2026, the review categorizes research trends into handwritten text recognition (HTR), optical character recognition (OCR), script classification, dataset development, and multimodal vision–language systems. The findings reveal a significant architectural shift from traditional segmentation-based CNN and RNN models toward transformer architectures and multimodal approaches. However, for Chagatai specifically, the primary obstacle is not the lack of advanced models but a critical scarcity of basic research infrastructure, including expert-verified transcriptions, annotation standards, and open benchmark datasets. Consequently, this article proposes a concrete development roadmap focusing on systematic digitization, expert annotation, transfer learning, and the creation of baseline models to enable reproducible evaluations.

Read PDF

Similar papers

Open access Aug 2026

Good Enough to Read?

Despite the imperfect output often produced by Arabic HTR, this research demonstrates that LLMs can effectively derive meaningful summaries and key insights, enabling rapid triage for scholars, enabling rapid triage for scholars.

Albrecht Hofheinz · 0 citations
Open access Aug 2026

Extraction of Handwritten and Printed Cyrillic Text from Documents: A Resource-Efficient Pipeline

This paper presents a comprehensive, resource-efficient pipeline for extracting printed and handwritten Cyrillic text that integrates advanced image preprocessing, YOLO-based document structure and table recognition, and a Permuted Autoregressive Sequence (PARSeq) model trained specifically for Bulgarian.

D. Halachev, Ivan Koychev · 0 citations
Open access Aug 2026

Large language models for OCR in cultural heritage: a comparative study on Slovene Folkloristic texts

This study provides the first systematic benchmark of LLM-based OCR for Slovene folkloristic materials, offering practical insights for cultural heritage digitization and highlighting the document-sensitivity of OCR performance.

O. Machidon, Jasmina Rejec, Domen Vres et al. · 0 citations
Open access Aug 2026

UniLipi: A Unified Multi-script OCR for Historical Indic Manuscripts

UniLipi, a unified multi-script OCR model for handwritten Indic manuscripts trained jointly across 13 Indic scripts within a single framework, serves as an effective foundational pretrained model and predicts script identity and per-line native character counts, supporting practical manuscript cataloging workflows.

Tathagata Ghosh, Sai Madhusudan Gunda, Simran Singh Sandral et al. · 0 citations

HIT. A hybrid OCR methodology for document analysis for historical documents

This work provides a validated, privacy-preserving, and locally deployable solution for the high-fidelity transcription of sensitive human rights archives through a hybrid methodology that combines domain-specific fine-tuning for text recognition models with a novel anchoring mechanism to ground VLM generation.

Cristobal Sebastian Vasquez Rosel · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.