This work provides a validated, privacy-preserving, and locally deployable solution for the high-fidelity transcription of sensitive human rights archives through a hybrid methodology that combines domain-specific fine-tuning for text recognition models with a novel anchoring mechanism to ground VLM generation.
Abstract
The transcription of historical documents from the Chilean dictatorship (1973–1990) is essential for the preservation of memory and the pursuit of justice. However, these archives present significant challenges due to severe physical degradation, noise, and typographic variability, which cause standard Optical Character Recognition (OCR) systems and modern Vision-Language Models (VLMs) to struggle, often resulting in hallucinations or low-fidelity outputs. This thesis addresses this problem by proposing HiT (Here is the Text), a hybrid methodology that combines domain-specific fine-tuning for text recognition models with a novel anchoring mechanism to ground VLM generation. The method consists of two stages. First, we introduce DHiSS and DHiSS+, the first large-scale word-level datasets for this domain, comprising over 110,000 and 185,000 curated images respectively. Second, we present the HiT anchoring pipeline, which leverages a text recognition model, fine-tuned on these datasets to inject high-confidence lexical and geometric cues into a VLM. Experimental results on a representative test set demonstrate that the proposed approach significantly outperforms both local and commercial cloud-based baselines. The best configuration (HiT-DHiSS+) achieves a Word Error Rate of 0.162, representing a 41.9% reduction compared to the unanchored baseline (0.279), while maintaining robustness across a wide range of confidence thresholds. This work provides a validated, privacy-preserving, and locally deployable solution for the high-fidelity transcription of sensitive human rights archives.
A scalable evaluation framework for simile understanding is proposed that includes a controlled simile dataset in which metaphorical vehicles are drawn from a predefined set of object-detectable categories and combined with diverse templates.
Luecheng Wang, Shintaro Ozaki, Hidetaka Kamigaito et al.· 0 citations
This work adapts PP-OCRv6, a recent compact text recognizer without strong language modeling, for historical line recognition and compares it with a conventional CRNN across generalized pretraining, domain-specific training, corpus-level fine-tuning, and manuscript-specific few-shot adaptation on multilingual Latin- an...
Experiments on a new approach for fine-grained evaluation demonstrate that this approach enhances a model’s ability to understand fine-grained differences.
Aozhu Chen, Hazel Doughty, Xirong Li et al.· 0 citations
In recent years, text-to-image (T2I) generation models have made substantial progress, particularly in visual realism and the expression of prompt semantics. However, a key difficulty remains: how to evaluate generated results automatically in a way that is both comprehensive and interpretable, while still being practi...
Optical Character Recognition (OCR) is a key component in the digitization of historical archives. Recently, Vision-Language Models (VLMs) have emerged as strong alternatives to traditional OCR systems, achieving state-of-the-art performance on standard benchmarks. However, their suitability for archival transcription...
Marina Gardella, Camilo Mariño, Diego Belzarena et al.· arXiv.org· 0 citations