OCRs for Corpus Extraction for the Maltese Language
Abstract
This paper presents the DocEng 2026 Competition on Maltese Optical Character Recognition (OCR). The competition challenged participants to develop OCR systems capable of accurately transcribing paragraph images extracted from Maltese-language PDF documents into single-line text suitable for corpus construction. As no annotated OCR training set was provided, participants were required to generate their own synthetic training data, while development and held-out test sets were supplied for validation and evaluation. Three teams participated, exploring approaches based on fine-tuned Tesseract models, transformer architectures, and ensemble methods. The results demonstrate that synthetic training data, combined with effective language-aware modelling and postprocessing, can produce highly accurate OCR systems for Maltese document transcription.