Skip to content
Open access

On Curating HTR Training Datasets for Romanian Language with use of Transcribathon Tool

Aug 2026 · International Journal of Digital Curation · 0 citations

TL;DR

The workflow applied to prepare datasets of Romanian historical documents for training language-specific HTR models or for enhancing the language coverage of general purpose models is presented and a preliminary evaluation of the effort required to create HTR training sets is presented.

Abstract

Fulltext digitisation of historical manuscripts is an important activity in Digital Archives. The impressive development of handwritten text recognition technology provides the required technical support for mass (fulltext) digitization of archival materials. The quality of the automatically recognised text has become high for a couple of intensively used languages like English, German, French, etc. However, this is not the case for under-represented languages such as Romanian, for which, according to our current level of knowledge, there is currently no HTR model openly available. While Transkribus provides more than 300 public models, none of them support the Romanian language in Latin script. There are two experimental models available in Transkribus (https://www.transkribus.org/), one for Romanian in Cyrillic script and one for the transitional script, both being trained on relatively small datasets. In this paper, we present the workflow applied to prepare datasets of Romanian historical documents for training language-specific HTR models or for enhancing the language coverage of general purpose models. The process for creating such datasets uses the functionality implemented within the Transcribathon.eu tool for combining pre-existing manual transcriptions with automatically extracted HTR text. A preliminary evaluation of the effort required to create HTR training sets and an assessment of their quality is presented in this paper, allowing us to draw conclusions on the proposed approach.

Read PDF

Similar papers

Preprint Sep 2026

Exploring In-Context Learning for Handwritten Text Recognition

Handwritten Text Recognition (HTR) systems have become an indispensable tool for the digitization of historical documents. Not only do they cut down time and cost, but they also allow democratizing access and processing of their contents by generating their transcripts. However, literature in HTR currently focuses most...

Eric Ayllon, Abel Gandia, Jorge Calvo-Zaragoza · 0 citations
Open access Aug 2026

Dataset Curation for Kalabari NMT System

A transferable curation framework for endangered languages, the first sizable Kalabari-English parallel corpus, and baseline experiments that reveal both the promise and the hallucination pitfalls of training on highly constrained, domain-specific data are contributed.

O. T. Olise · 0 citations
Book Open access Aug 2026

OCRs for Corpus Extraction for the Maltese Language

This paper presents the DocEng 2026 Competition on Maltese Optical Character Recognition (OCR). The competition challenged participants to develop OCR systems capable of accurately transcribing paragraph images extracted from Maltese-language PDF documents into single-line text suitable for corpus construction. As no a...

Marc Tanti, Stefania Cristina, Alexandra Bonnici · 0 citations
#machine learning Preprint Sep 2026

Measuring Annotation Efficiency for Handwritten Devanagari Recognition: Sample-Complexity Curves for Four Pretraining Regimes

The scarcity in this study is constructed by subsampling a large corpus by changing only the number of real transcribed words used for fine-tuning across nine budgets from 10 to 4,000 and four initialisation regimes, with six seeds at every point.

Manglesh Kumar Pandey, S. Banshal · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.