Skip to content
Open access

Character-Based Arabic Offline Handwritten Text Recognition Using Faster R-CNN

Aug 2026 · Applied Sciences · 0 citations · 24 references

TL;DR

These findings demonstrate that explicit character localization provides a robust, data-efficient alternative for Arabic handwritten text recognition in low-resource settings.

Abstract

Offline handwritten word recognition has progressed from whole-word classification to sequence transcription, yet many systems depend on large annotated corpora and exploit lexical regularities over explicit character evidence. This paper presents an alternative formulation for Arabic offline handwritten word recognition, treating characters as spatial objects detected via a Faster Region-Based Convolutional Neural Network rather than symbols generated by a one-dimensional decoder. We construct and release a character-level annotated subset of 2153 handwritten word images from a standard Arabic benchmark, exporting matched detection, sequence, and word-class labels. We also introduce an open-source subword exchange toolkit that creates a controlled structural-generalization benchmark by swapping subwords while preserving handwriting style. Experiments compare the proposed detector against whole-word and sequence-based baselines on both the original held-out split and the perturbed benchmark. Results show sequence models degrade sharply under structural recombination, whereas the proposed detector remains stable, achieving a 26.56% character error rate and 70.0% word accuracy on the perturbed benchmark. These findings demonstrate that explicit character localization provides a robust, data-efficient alternative for Arabic handwritten text recognition in low-resource settings.

Read PDF

Similar papers

Open access Aug 2026

Improving Right to Left Cursive Handwritten Text Recognition in Historical Manuscripts Using Learnable Edge Features and Channel Attention

An edge-aware line-level HTR framework that extends a CNN-Transformer baseline with a learnable edge-extraction channel and Squeeze-and-Excitation channel attention and shows that combining learnable structural cues with channel-wise attention has improved robustness for degradation-prone historical manuscript collecti...

Bilal Abdulrahman, Farhan Mohamed · 0 citations
Open access 2026

Handwritten Word Recognition for Low-Resource Languages: A CRNN-CTC Framework for Kirundi

Although exact word-level recognition remained difficult because of the extremely limited dataset size, the proposed framework successfully learned meaningful sequential patterns and produced increasingly structured Kirundi-like predictions.

Niyifasha Patrick · 0 citations
#computer vision Preprint Aug 2026

Towards a Joint Khmer Text Recognition and Word Segmentation

Experimental results show that the proposed model can not only recognize characters in document images but also locate word boundaries, removing the need for an extra word segmentation step in a conventional sequential pipeline.

Marry Kong, Rina Buoy, Sovisal Chenda et al. · 0 citations

Offline Bengali Writer Verification by PDF-CNN and Siamese Net

This paper deals with offline writer verification on complex handwriting patterns, where handcrafted feature PDFs are hybridized with auto-derived CNN features and fed into a Siamese neural network for writer verification.

Chandranath Adak, SimoneMarinai, BidyutB.Chaudhuri et al. · 0 citations
Preprint Aug 2026

DTRNet: Dual Text-Radical Decoding for Handwritten Chinese Text Recognition with Faked Character Detection

In K-12 educational scenarios, handwritten Chinese text recognition should not only transcribe student writing, but also detect faked characters. However, existing recognition models are usually confined to a predefined set of normal characters and therefore cannot explicitly identify faked characters. Existing detecti...

Runrui Li, Lin Zhu, Hua Huang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.