A lightweight framework optimized for efficient performance on devices with severely limited computational capacity, aimed at extracting coherent sentence-level segments from document images and demonstrating strong classification performance for distinguishing handwritten and printed sentence-level text images.
Abstract
With the increasing demand for reusing paper documents in educational and office settings, accurate segmentation of handwritten and printed text has become a crucial step in document digitization. Although numerous deep learning models have been developed for this task, their high computational cost limits deployment on resource-constrained edge devices. To address this challenge, we present a lightweight framework optimized for efficient performance on devices with severely limited computational capacity. Our approach begins with the Sentence-level Connected Component Segmentation algorithm, aimed at extracting coherent sentence-level segments from document images. We then design a novel Region-aware Handwriting Descriptor (RHD) to capture the intrinsic variability of human handwriting at the sentence level. A simple conventional classifier can then be seamlessly integrated with our designed descriptor, demonstrating strong classification performance for distinguishing handwritten and printed sentence-level text images, highlighting that the proposed descriptor is agnostic to the choice of classifier. Extensive experiments are performed on our self-constructed Multilingual High-Quality Annotated Dataset for Handwritten and Printed Text Segmentation (MAD-HPTS) and a public benchmark PHD-AS, and the experimental results demonstrate that the proposed framework outperforms current state-of-the-art methods in both accuracy and computational efficiency. On MAD-HPTS, our method sacrifices only 1.4% accuracy compared to the leading deep neural network baseline, yet achieves more than 8 times speedup in inference, making it well-suited for lightweight deployment.
These findings demonstrate that explicit character localization provides a robust, data-efficient alternative for Arabic handwritten text recognition in low-resource settings.
Sofiane Medjram, Ruwaidah Saud Alnejaidi· Applied Sciences· 0 citations
This paper presents a comprehensive, resource-efficient pipeline for extracting printed and handwritten Cyrillic text that integrates advanced image preprocessing, YOLO-based document structure and table recognition, and a Permuted Autoregressive Sequence (PARSeq) model trained specifically for Bulgarian.
D. Halachev, Ivan Koychev· Computational Linguistics in...· 0 citations
Providing precise character-level feedback on handwritten text requires systems capable of localizing individual characters. However, existing handwriting datasets typically lack character-level annotations, limiting the development of localization models. In this work, we introduce the first large-scale English data...
S. Yasin, T. Zesch· International Journal on Doc...· 0 citations
An edge-aware line-level HTR framework that extends a CNN-Transformer baseline with a learnable edge-extraction channel and Squeeze-and-Excitation channel attention and shows that combining learnable structural cues with channel-wise attention has improved robustness for degradation-prone historical manuscript collecti...
Bilal Abdulrahman, Farhan Mohamed· Journal of Human Centered Te...· 0 citations
Experimental results show that the proposed model can not only recognize characters in document images but also locate word boundaries, removing the need for an extra word segmentation step in a conventional sequential pipeline.
Marry Kong, Rina Buoy, Sovisal Chenda et al.· 0 citations
University lecturers often annotate printed slides during class with handwritten notes and mathematical derivations. Generic OCR skips this ink or returns unrelated strings, and no public dataset fits the task: German handwriting corpora hold prose or historical script, handwritten-mathematics corpora are language-neut...
N. Ragavenderan, Judith Jakob, S. Manonmani· IEEE Access· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.