Skip to content

IndicPage-OCR: Robust Low-Resource Adaptation for Multi-Script Indic Handwritten Page Recognition

· 0 citations · 18 references

TL;DR

Experiments show substantial reductions in Word Error Rate (WER) and Character Error Rate (CER), narrowing the performance gap between commercial and freely deployable OCR systems by approximately 80% in WER and over 90% in CER, while consistently outperforming general-purpose foundation-style baselines.

View source

Similar papers

Open access Aug 2026

UniLipi: A Unified Multi-script OCR for Historical Indic Manuscripts

UniLipi, a unified multi-script OCR model for handwritten Indic manuscripts trained jointly across 13 Indic scripts within a single framework, serves as an effective foundational pretrained model and predicts script identity and per-line native character counts, supporting practical manuscript cataloging workflows.

Tathagata Ghosh, Sai Madhusudan Gunda, Simran Singh Sandral et al. · 0 citations
Preprint Sep 2026

Cross-Domain Few-Shot Writer Adaptation for Real-World Handwritten Mathematical Expression Recognition

Handwritten mathematical expression recognition (HMER) refers to the task of recognizing and converting handwritten mathematics into a parsable markup language, usually LaTeX. No current state-of-the-art-competitive system adjusts to the way a specific person writes, and the domain gap between training images (usually...

P. Silva, Lorenz Bernard Marqueses, J. Ilao · 0 citations
Preprint Sep 2026

All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts

This work constructs TextMuSS-10M, a large-scale synthetic scene text dataset spanning 10 scripts and 229 languages and proposes ScriptMoE, a script-aware Mixture-of-Experts (MoE) architecture that achieves the highest accuracy and is simpler than per-language experts, lighter than VLMs, and more accurate than both.

Xing-Song Ye, Yong-Kun Du, Jia-Xin Zhang et al. · 1 citation
Aug 2026

TamilLite-Gan: A lightweight graph attention network for handwritten Tamil character recognition using enhanced single shot optimization

Handwritten character recognition plays a crucial role in optical character recognition systems, particularly for low-resource and structurally complex scripts such as Tamil. Despite significant advances in deep learning, accurate recognition of handwritten Tamil characters remains challenging due to large variations i...

K. Manoj, M. Iyapparaja · 0 citations
Open access 2026

Diffusion-Enhanced NAT–BART Vision Language Transformer for Unified Urdu Word Recognition

This is the first study to introduce both a real handwritten Urdu word dataset and a diffusion-generated synthetic dataset, and develops a unified word recognition model trained jointly on handwritten and printed Urdu word data, leading to improved recognition robustness and performance.

Wahid Hussain, Shahbaz Hassan, I. Hassan et al. · 0 citations
Open access 2026

Multi-Scale Transformer-Based Lexicon-Guided Handwritten Text Recognition Using Adaptive Feature Fusion

This paper proposes a Multi-Scale Transformer-Based Lexicon-Guided HTR framework built around an Adaptive Feature Fusion (AFF) mechanism, which trains a Transformer encoder with multi-head self-attention that models long-range context far more effectively than bidirectional recurrent layers.

Lalita Kumari · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.