Aug 2026· International Journal on Document Analysis and Recognition· 0 citations· 16 references
TL;DR
This study provides the first systematic benchmark of LLM-based OCR for Slovene folkloristic materials, offering practical insights for cultural heritage digitization and highlighting the document-sensitivity of OCR performance.
Abstract
The digitization of historical and folkloristic texts presents significant challenges for optical character recognition (OCR), particularly when documents contain complex layouts, embedded illustrations, irregular typography, or non-standard language. This study provides a systematic evaluation of six OCR approaches on two Slovene-language heritage collections: typewritten folklore manuscripts with uniform formatting, and visually heterogeneous issues of the children’s magazine
Ciciban
. The methods compared include Tesseract, Tesseract with GPT 5.2 post-processing, GPT 5.2 direct transcription, LLaMA 4 Maverick, Nanonets OCR-3, and Qwen-VL-OCR. Performance was assessed using character error rate, word error rate, and complementary sequence-based metrics against manually aligned ground truth. Results indicate that direct multimodal and document-oriented systems achieve the strongest accuracy on typewritten texts, while performance on
Ciciban
is more sensitive to layout structure. These findings highlight the document-sensitivity of OCR performance and point to the need for adaptive, content-aware pipelines that dynamically integrate multiple OCR strategies. To our knowledge, the study provides the first systematic benchmark of LLM-based OCR for Slovene folkloristic materials, offering practical insights for cultural heritage digitization.
Purpose: The digitization of historical documents presents fundamental challenges for modern information retrieval and Artificial Intelligence (AI) systems. Optical character recognition (OCR) errors in source corpora propagate through retrieval-augmented generation (RAG) pipelines, compromising the factual accuracy of...
Marina Gómez Rey, Patricia Callejo, M. Muñoz-Organero et al.· 0 citations
This work introduces BANGLAWILD, a benchmark of 2,535 Bengali scene text images, each paired with a verbatim gold transcription, two categorical axes, four diagnostic attributes, and an orthographically standard form where the in-image text deviates from canonical spelling.
Sadab Shiper, Tawsif Tashwar Dipto, M. Inzamam et al.· 0 citations
This work provides a validated, privacy-preserving, and locally deployable solution for the high-fidelity transcription of sensitive human rights archives through a hybrid methodology that combines domain-specific fine-tuning for text recognition models with a novel anchoring mechanism to ground VLM generation.
The workflow applied to prepare datasets of Romanian historical documents for training language-specific HTR models or for enhancing the language coverage of general purpose models is presented and a preliminary evaluation of the effort required to create HTR training sets is presented.
S. Gordea, George Cristian Cotea, Frank Drauschke et al.· International Journal of Dig...· 0 citations
OmniHandwritingOCR provides a challenging testbed for diagnosing language, content, structural, and visual-grounding failure modes of multimodal models in handwritten OCR scenarios.
Zinuo Guo, Min Zhang, Bo Jiang· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.