Skip to content
Book Open access

OCRs for Corpus Extraction for the Maltese Language

Aug 2026 · Proceedings of the 2026 ACM Symposium on Document Engineering · 0 citations · 8 references

Abstract

This paper presents the DocEng 2026 Competition on Maltese Optical Character Recognition (OCR). The competition challenged participants to develop OCR systems capable of accurately transcribing paragraph images extracted from Maltese-language PDF documents into single-line text suitable for corpus construction. As no annotated OCR training set was provided, participants were required to generate their own synthetic training data, while development and held-out test sets were supplied for validation and evaluation. Three teams participated, exploring approaches based on fine-tuned Tesseract models, transformer architectures, and ensemble methods. The results demonstrate that synthetic training data, combined with effective language-aware modelling and postprocessing, can produce highly accurate OCR systems for Maltese document transcription.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.