Skip to content
Open access

Extraction of Handwritten and Printed Cyrillic Text from Documents: A Resource-Efficient Pipeline

Aug 2026 · Computational Linguistics in Bulgaria · 0 citations

TL;DR

This paper presents a comprehensive, resource-efficient pipeline for extracting printed and handwritten Cyrillic text that integrates advanced image preprocessing, YOLO-based document structure and table recognition, and a Permuted Autoregressive Sequence (PARSeq) model trained specifically for Bulgarian.

Abstract

The digitisation of historical, administrative, and personal documents in Bulgarian faces considerable challenges due to the lack of robust Optical Character Recognition (OCR) and Handwritten Text Recognition (HTR) systems tailored for the Cyrillic alphabet. While modern Vision-Language Models (VLMs) and large transformer-based architectures achieve state-ofthe-art results, their performance and resource efficiency on low-resource languages remain prohibitive for decentralised, privacy-preserving applications. In this paper, we present a comprehensive, resource-efficient pipeline for extracting printed and handwritten Cyrillic text. Our system integrates advanced image preprocessing, YOLO-based document structure and table recognition, and a Permuted Autoregressive Sequence (PARSeq) model trained specifically for Bulgarian. We generated custom synthetic Bulgarian cursive datasets to mitigate the severe lack of real-world training data. Our evaluation indicates that the specialised PARSeq model outperforms traditional OCR tools such as Tesseract and EasyOCR on our custom degraded printed test set, and provides a practical, resource-efficient baseline for handwriting recognition compared to a modern local VLM (Qwen3-VL-4B). Finally, we discuss the discrepancy between synthetic and real handwritten data, highlighting the urgent need for a standardised, annotated Bulgarian HTR dataset.

Read PDF

Similar papers

IndicPage-OCR: Robust Low-Resource Adaptation for Multi-Script Indic Handwritten Page Recognition

Experiments show substantial reductions in Word Error Rate (WER) and Character Error Rate (CER), narrowing the performance gap between commercial and freely deployable OCR systems by approximately 80% in WER and over 90% in CER, while consistently outperforming general-purpose foundation-style baselines.

Shaon Bhattacharyya, Ajoy Mondal, C. V. Jawahar et al. · 0 citations
Open access 2026

Handwritten Word Recognition for Low-Resource Languages: A CRNN-CTC Framework for Kirundi

Although exact word-level recognition remained difficult because of the extremely limited dataset size, the proposed framework successfully learned meaningful sequential patterns and produced increasingly structured Kirundi-like predictions.

Niyifasha Patrick · 0 citations
Open access Aug 2026

Character-Based Arabic Offline Handwritten Text Recognition Using Faster R-CNN

These findings demonstrate that explicit character localization provides a robust, data-efficient alternative for Arabic handwritten text recognition in low-resource settings.

Sofiane Medjram, Ruwaidah Saud Alnejaidi · 0 citations
Open access Aug 2026

Improving Right to Left Cursive Handwritten Text Recognition in Historical Manuscripts Using Learnable Edge Features and Channel Attention

An edge-aware line-level HTR framework that extends a CNN-Transformer baseline with a learnable edge-extraction channel and Squeeze-and-Excitation channel attention and shows that combining learnable structural cues with channel-wise attention has improved robustness for degradation-prone historical manuscript collecti...

Bilal Abdulrahman, Farhan Mohamed · 0 citations
Open access 2026

Typeset Replacement of Handwritten Text and Mathematics on Lecture Slides Using Vision-Language Models

University lecturers often annotate printed slides during class with handwritten notes and mathematical derivations. Generic OCR skips this ink or returns unrelated strings, and no public dataset fits the task: German handwriting corpora hold prose or historical script, handwritten-mathematics corpora are language-neut...

N. Ragavenderan, Judith Jakob, S. Manonmani · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.