Skip to content
Open access

To Overfit or Not to Overfit? An Evaluation of HTR Workflow on 17Th-18Th Century French Corpus

2026 · International Conference on Language Resources and Evaluation · pp. 1009-1016 · 0 citations · 30 references
Computer Science

TL;DR

An evaluation of general Handwritten Text Recognition models applied to 17th and 18th century corpus written in modern French and the fine-tuning of the models shows improved transcription accuracy and reduced processing time.

Abstract

This paper presents the results of an evaluation of general Handwritten Text Recognition (HTR) models applied to 17th and 18th century corpus written in modern French and the fine-tuning of the models. Our aim was to transcribe a corpus from this period using existing pre-trained models and to assess their performance on such data. While these general models offer a large linguistic coverage, our results demonstrate they are often insufficiently adapted to the specific handwriting nuances and orthographic inconsistencies of early modern French. To improve the results, we fine-tuned a base model to develop a specialized version trained on our dataset. Although the model still encountered difficulties due to highly variable handwriting styles, it significantly improved transcription accuracy and reduced processing time. Following this step, we used a semi-automatic post-correction tool to address remaining errors and integrated Named Entity Recognition (NER) steps for automated TEI-XML encoding. This paper discusses the evaluation results of both the HTR and NER models, and how the overfitting allows to get better transcriptions on a specific corpus.

Read PDF

Similar papers

Preprint Sep 2026

Exploring In-Context Learning for Handwritten Text Recognition

Handwritten Text Recognition (HTR) systems have become an indispensable tool for the digitization of historical documents. Not only do they cut down time and cost, but they also allow democratizing access and processing of their contents by generating their transcripts. However, literature in HTR currently focuses most...

Eric Ayllon, Abel Gandia, Jorge Calvo-Zaragoza · 0 citations
Review Open access Sep 2026

AI-ASSISTED PIPELINE FOR SEMI-AUTOMATIC CONSTRUCTION OF HTR CORPORA FROM HISTORICAL DOCUMENTS

This paper presents an improved software application for the semi-automatic annotation of handwritten words from historical documents using a CRNN deep learning model. The proposed approach integrates an artificial intelligence module directly into the annotation process, automatically generating a preliminary transcri...

A. Ivasechko, Khrystyna Lipianina-Honcharenko · 0 citations
Open access Aug 2026

Good Enough to Read?

Despite the imperfect output often produced by Arabic HTR, this research demonstrates that LLMs can effectively derive meaningful summaries and key insights, enabling rapid triage for scholars, enabling rapid triage for scholars.

Albrecht Hofheinz · 0 citations
Open access Aug 2026

On Curating HTR Training Datasets for Romanian Language with use of Transcribathon Tool

The workflow applied to prepare datasets of Romanian historical documents for training language-specific HTR models or for enhancing the language coverage of general purpose models is presented and a preliminary evaluation of the effort required to create HTR training sets is presented.

S. Gordea, George Cristian Cotea, Frank Drauschke et al. · 0 citations
#machine learning Preprint Sep 2026

Measuring Annotation Efficiency for Handwritten Devanagari Recognition: Sample-Complexity Curves for Four Pretraining Regimes

The scarcity in this study is constructed by subsampling a large corpus by changing only the number of real transcribed words used for fine-tuning across nine budgets from 10 to 4,000 and four initialisation regimes, with six seeds at every point.

Manglesh Kumar Pandey, S. Banshal · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.