2026· International Conference on Language Resources and Evaluation· pp. 1009-1016· 0 citations· 30 references
Computer Science
TL;DR
An evaluation of general Handwritten Text Recognition models applied to 17th and 18th century corpus written in modern French and the fine-tuning of the models shows improved transcription accuracy and reduced processing time.
Abstract
This paper presents the results of an evaluation of general Handwritten Text Recognition (HTR) models applied to 17th and 18th century corpus written in modern French and the fine-tuning of the models. Our aim was to transcribe a corpus from this period using existing pre-trained models and to assess their performance on such data. While these general models offer a large linguistic coverage, our results demonstrate they are often insufficiently adapted to the specific handwriting nuances and orthographic inconsistencies of early modern French. To improve the results, we fine-tuned a base model to develop a specialized version trained on our dataset. Although the model still encountered difficulties due to highly variable handwriting styles, it significantly improved transcription accuracy and reduced processing time. Following this step, we used a semi-automatic post-correction tool to address remaining errors and integrated Named Entity Recognition (NER) steps for automated TEI-XML encoding. This paper discusses the evaluation results of both the HTR and NER models, and how the overfitting allows to get better transcriptions on a specific corpus.
Handwritten Text Recognition (HTR) systems have become an indispensable tool for the digitization of historical documents. Not only do they cut down time and cost, but they also allow democratizing access and processing of their contents by generating their transcripts. However, literature in HTR currently focuses most...
Eric Ayllon, Abel Gandia, Jorge Calvo-Zaragoza· 0 citations
This paper presents an improved software application for the semi-automatic annotation of handwritten words from historical documents using a CRNN deep learning model. The proposed approach integrates an artificial intelligence module directly into the annotation process, automatically generating a preliminary transcri...
A. Ivasechko, Khrystyna Lipianina-Honcharenko· Computer Systems and Informa...· 0 citations
Despite the imperfect output often produced by Arabic HTR, this research demonstrates that LLMs can effectively derive meaningful summaries and key insights, enabling rapid triage for scholars, enabling rapid triage for scholars.
Albrecht Hofheinz· Journal of Arabic and Islami...· 0 citations
The workflow applied to prepare datasets of Romanian historical documents for training language-specific HTR models or for enhancing the language coverage of general purpose models is presented and a preliminary evaluation of the effort required to create HTR training sets is presented.
S. Gordea, George Cristian Cotea, Frank Drauschke et al.· International Journal of Dig...· 0 citations
OmniHandwritingOCR provides a challenging testbed for diagnosing language, content, structural, and visual-grounding failure modes of multimodal models in handwritten OCR scenarios.
The scarcity in this study is constructed by subsampling a large corpus by changing only the number of real transcribed words used for fine-tuning across nine budgets from 10 to 4,000 and four initialisation regimes, with six seeds at every point.
Manglesh Kumar Pandey, S. Banshal· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.