Aug 2026· Computational Linguistics in Bulgaria· 0 citations
TL;DR
This paper presents a comprehensive, resource-efficient pipeline for extracting printed and handwritten Cyrillic text that integrates advanced image preprocessing, YOLO-based document structure and table recognition, and a Permuted Autoregressive Sequence (PARSeq) model trained specifically for Bulgarian.
Abstract
The digitisation of historical, administrative, and personal documents in Bulgarian faces considerable challenges due to the lack of robust Optical Character Recognition (OCR) and Handwritten Text Recognition (HTR) systems tailored for the Cyrillic alphabet. While modern Vision-Language Models (VLMs) and large transformer-based architectures achieve state-ofthe-art results, their performance and resource efficiency on low-resource languages remain prohibitive for decentralised, privacy-preserving applications. In this paper, we present a comprehensive, resource-efficient pipeline for extracting printed and handwritten Cyrillic text. Our system integrates advanced image preprocessing, YOLO-based document structure and table recognition, and a Permuted Autoregressive Sequence (PARSeq) model trained specifically for Bulgarian. We generated custom synthetic Bulgarian cursive datasets to mitigate the severe lack of real-world training data. Our evaluation indicates that the specialised PARSeq model outperforms traditional OCR tools such as Tesseract and EasyOCR on our custom degraded printed test set, and provides a practical, resource-efficient baseline for handwriting recognition compared to a modern local VLM (Qwen3-VL-4B). Finally, we discuss the discrepancy between synthetic and real handwritten data, highlighting the urgent need for a standardised, annotated Bulgarian HTR dataset.
Experiments show substantial reductions in Word Error Rate (WER) and Character Error Rate (CER), narrowing the performance gap between commercial and freely deployable OCR systems by approximately 80% in WER and over 90% in CER, while consistently outperforming general-purpose foundation-style baselines.
Shaon Bhattacharyya, Ajoy Mondal, C. V. Jawahar et al.· 0 citations
Although exact word-level recognition remained difficult because of the extremely limited dataset size, the proposed framework successfully learned meaningful sequential patterns and produced increasingly structured Kirundi-like predictions.
Niyifasha Patrick· International journal of re...· 0 citations
These findings demonstrate that explicit character localization provides a robust, data-efficient alternative for Arabic handwritten text recognition in low-resource settings.
Sofiane Medjram, Ruwaidah Saud Alnejaidi· Applied Sciences· 0 citations
An edge-aware line-level HTR framework that extends a CNN-Transformer baseline with a learnable edge-extraction channel and Squeeze-and-Excitation channel attention and shows that combining learnable structural cues with channel-wise attention has improved robustness for degradation-prone historical manuscript collecti...
Bilal Abdulrahman, Farhan Mohamed· Journal of Human Centered Te...· 0 citations
OmniHandwritingOCR provides a challenging testbed for diagnosing language, content, structural, and visual-grounding failure modes of multimodal models in handwritten OCR scenarios.
University lecturers often annotate printed slides during class with handwritten notes and mathematical derivations. Generic OCR skips this ink or returns unrelated strings, and no public dataset fits the task: German handwriting corpora hold prose or historical script, handwritten-mathematics corpora are language-neut...
N. Ragavenderan, Judith Jakob, S. Manonmani· IEEE Access· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.