Skip to content

HIT. A hybrid OCR methodology for document analysis for historical documents

TL;DR

This work provides a validated, privacy-preserving, and locally deployable solution for the high-fidelity transcription of sensitive human rights archives through a hybrid methodology that combines domain-specific fine-tuning for text recognition models with a novel anchoring mechanism to ground VLM generation.

Abstract

The transcription of historical documents from the Chilean dictatorship (1973–1990) is essential for the preservation of memory and the pursuit of justice. However, these archives present significant challenges due to severe physical degradation, noise, and typographic variability, which cause standard Optical Character Recognition (OCR) systems and modern Vision-Language Models (VLMs) to struggle, often resulting in hallucinations or low-fidelity outputs. This thesis addresses this problem by proposing HiT (Here is the Text), a hybrid methodology that combines domain-specific fine-tuning for text recognition models with a novel anchoring mechanism to ground VLM generation. The method consists of two stages. First, we introduce DHiSS and DHiSS+, the first large-scale word-level datasets for this domain, comprising over 110,000 and 185,000 curated images respectively. Second, we present the HiT anchoring pipeline, which leverages a text recognition model, fine-tuned on these datasets to inject high-confidence lexical and geometric cues into a VLM. Experimental results on a representative test set demonstrate that the proposed approach significantly outperforms both local and commercial cloud-based baselines. The best configuration (HiT-DHiSS+) achieves a Word Error Rate of 0.162, representing a 41.9% reduction compared to the unanchored baseline (0.279), while maintaining robustness across a wide range of confidence thresholds. This work provides a validated, privacy-preserving, and locally deployable solution for the high-fidelity transcription of sensitive human rights archives.

View source

Similar papers

Preprint Aug 2026

Simile Understanding in Text-to-Image Models: An Evaluation Framework

A scalable evaluation framework for simile understanding is proposed that includes a controlled simile dataset in which metaphorical vehicles are drawn from a predefined set of object-detectable categories and combined with diverse templates.

Luecheng Wang, Shintaro Ozaki, Hidetaka Kamigaito et al. · 0 citations
Preprint Sep 2026

A Free Lunch? Adapting PP-OCRv6 for Historical Text Recognition

This work adapts PP-OCRv6, a recent compact text recognizer without strong language modeling, for historical line recognition and compares it with a conventional CRNN across generalized pretraining, domain-specific training, corpus-level fine-tuning, and manuscript-specific few-shot adaptation on multilingual Latin- an...

Benjamin Kiessling · 0 citations
Open access Aug 2026

Structured Creative Evaluation for Text-to-Image Generative AI Models

In recent years, text-to-image (T2I) generation models have made substantial progress, particularly in visual realism and the expression of prompt semantics. However, a key difficulty remains: how to evaluate generated results automatically in a way that is both comprehensive and interpretable, while still being practi...

W.-C. Ma, Q. Zhang · 0 citations
#computer vision Jul 2026

When Low CER is Not Enough: An Analysis of Hallucinations in Vision-Language OCR Systems on Historical Uruguayan Documents

Optical Character Recognition (OCR) is a key component in the digitization of historical archives. Recently, Vision-Language Models (VLMs) have emerged as strong alternatives to traditional OCR systems, achieving state-of-the-art performance on standard benchmarks. However, their suitability for archival transcription...

Marina Gardella, Camilo Mariño, Diego Belzarena et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.