Skip to content

Measuring Annotation Efficiency for Handwritten Devanagari Recognition: Sample-Complexity Curves for Four Pretraining Regimes

Sep 2026 · 0 citations · 22 references
Computer Science

TL;DR

The scarcity in this study is constructed by subsampling a large corpus by changing only the number of real transcribed words used for fine-tuning across nine budgets from 10 to 4,000 and four initialisation regimes, with six seeds at every point.

Abstract

To train handwritten text recognition systems we need word images and their corresponding transcriptions, and these transcriptions are produced manually. For a script that can be read by only a small number of specialists, this manual transcription is a limitation, because the trained models are supposed to save the time of those same specialists. A relevant question therefore arises: how many transcriptions are needed before a recogniser becomes useful, and how much of that cost can pretraining remove? In this study the answer is measured directly for handwritten Devanagari. We keep the recogniser, optimiser and evaluation protocol the same and change only the number of real transcribed words used for fine-tuning across nine budgets from 10 to 4,000 and four initialisation regimes, with six seeds at every point. The resulting curves are then converted into annotation-equivalent terms. A CER of 0.50 is reached by supervised synthetic pretraining using only 81 transcribed words, whereas random initialisation requires 355, which gives a label multiplier of 4.40 [3.56, 4.99]. There is a zero-shot reference point as well: with no real transcribed words at all, this pretraining is worth about 136 of them. This advantage gets smaller as the target accuracy improves, and at the most demanding target we measure, it cannot be distinguished from no saving at all. A fourth arm in which only the encoder is transferred separates the effect of the pretraining method from that of transfer scope, and masked image modelling is observed to transfer negatively over a bounded range of budgets. We emphasise that the scarcity in this study is constructed by subsampling a large corpus.

View source

Similar papers

Preprint Sep 2026

Exploring In-Context Learning for Handwritten Text Recognition

Handwritten Text Recognition (HTR) systems have become an indispensable tool for the digitization of historical documents. Not only do they cut down time and cost, but they also allow democratizing access and processing of their contents by generating their transcripts. However, literature in HTR currently focuses most...

Eric Ayllon, Abel Gandia, Jorge Calvo-Zaragoza · 0 citations
Preprint Sep 2026

Cross-Domain Few-Shot Writer Adaptation for Real-World Handwritten Mathematical Expression Recognition

Handwritten mathematical expression recognition (HMER) refers to the task of recognizing and converting handwritten mathematics into a parsable markup language, usually LaTeX. No current state-of-the-art-competitive system adjusts to the way a specific person writes, and the domain gap between training images (usually...

P. Silva, Lorenz Bernard Marqueses, J. Ilao · 0 citations
Open access 2026

Typeset Replacement of Handwritten Text and Mathematics on Lecture Slides Using Vision-Language Models

This work asks whether a general vision-language model (VLM) can outperform purpose-built OCR and handwritten-mathematics recognisers in this data-scarce setting, and drives a pipeline that typesets the recognised annotations back onto the slide.

N. Ragavenderan, Judith Jakob, S. Manonmani · 0 citations
#machine learning Preprint Sep 2026

ExpertHTR: Unified Handwritten Text Recognition with Multi-Task Learning and Sparse Mixture-of-Experts

Experiments on seven heterogeneous handwriting benchmarks show that complementary supervision consistently improves training with page transcription alone, while joint multi-source training provides further gains on most datasets, and the proposed sparse expert model further improves the dense baseline on six of the se...

Nam Hoai Dang, H. Nguyen, Quang Huu Hieu et al. · 0 citations
#natural language process... Preprint Sep 2026

The Hidden Cost of Digits: Number Normalization and WER in ASR Systems

This work performs experiments on VoxPopuli and The Polish Parliamentary speech datasets and estimates word error rate (WER) differences for different text normalization approaches, showing that the difference due to the lack of number normalization in WER may be substantial.

Stanisław Kacprzak, Mieszko Fraś · 0 citations

Related blog posts

Microsoft Research Blog Aug 11, 2026

Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

Radiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning, calibrated predictions, and measurement-based tools for chest X-ray interpretation. The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research.

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.