Skip to content
Open access

Large language models for OCR in cultural heritage: a comparative study on Slovene Folkloristic texts

Aug 2026 · International Journal on Document Analysis and Recognition · 0 citations · 16 references

TL;DR

This study provides the first systematic benchmark of LLM-based OCR for Slovene folkloristic materials, offering practical insights for cultural heritage digitization and highlighting the document-sensitivity of OCR performance.

Abstract

The digitization of historical and folkloristic texts presents significant challenges for optical character recognition (OCR), particularly when documents contain complex layouts, embedded illustrations, irregular typography, or non-standard language. This study provides a systematic evaluation of six OCR approaches on two Slovene-language heritage collections: typewritten folklore manuscripts with uniform formatting, and visually heterogeneous issues of the children’s magazine Ciciban . The methods compared include Tesseract, Tesseract with GPT 5.2 post-processing, GPT 5.2 direct transcription, LLaMA 4 Maverick, Nanonets OCR-3, and Qwen-VL-OCR. Performance was assessed using character error rate, word error rate, and complementary sequence-based metrics against manually aligned ground truth. Results indicate that direct multimodal and document-oriented systems achieve the strongest accuracy on typewritten texts, while performance on Ciciban is more sensitive to layout structure. These findings highlight the document-sensitivity of OCR performance and point to the need for adaptive, content-aware pipelines that dynamically integrate multiple OCR strategies. To our knowledge, the study provides the first systematic benchmark of LLM-based OCR for Slovene folkloristic materials, offering practical insights for cultural heritage digitization.

Read PDF

Similar papers

Preprint Aug 2026

A Comparative Evaluation of Digitization Pipelines for Historiographical Sources

Purpose: The digitization of historical documents presents fundamental challenges for modern information retrieval and Artificial Intelligence (AI) systems. Optical character recognition (OCR) errors in source corpora propagate through retrieval-augmented generation (RAG) pipelines, compromising the factual accuracy of...

Marina Gómez Rey, Patricia Callejo, M. Muñoz-Organero et al. · 0 citations
Preprint Aug 2026

BanglaWild: An In-the-Wild Bengali Scene Text Recognition Benchmark for OCR and Vision-Language Models

This work introduces BANGLAWILD, a benchmark of 2,535 Bengali scene text images, each paired with a verbatim gold transcription, two categorical axes, four diagnostic attributes, and an orthographically standard form where the in-image text deviates from canonical spelling.

Sadab Shiper, Tawsif Tashwar Dipto, M. Inzamam et al. · 0 citations

HIT. A hybrid OCR methodology for document analysis for historical documents

This work provides a validated, privacy-preserving, and locally deployable solution for the high-fidelity transcription of sensitive human rights archives through a hybrid methodology that combines domain-specific fine-tuning for text recognition models with a novel anchoring mechanism to ground VLM generation.

Cristobal Sebastian Vasquez Rosel · 0 citations
Open access Aug 2026

On Curating HTR Training Datasets for Romanian Language with use of Transcribathon Tool

The workflow applied to prepare datasets of Romanian historical documents for training language-specific HTR models or for enhancing the language coverage of general purpose models is presented and a preliminary evaluation of the effort required to create HTR training sets is presented.

S. Gordea, George Cristian Cotea, Frank Drauschke et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.