Skip to content
Open access

A Confidence-Aware Hybrid OCR and Visual Retrieval Framework for Arabic Document Images

2026 · International Journal of Applied Science and Research · 0 citations

Abstract

The high rate of digitization of Arabic institutional archives has created masses of page images, which are hard to search where good textual metadata are unavailable. Lexical indexing can be supported by Optical Character Recognition (OCR), but its reliability suf ers from blur, skew, low contrast, compression artifacts, complex layout, and Arabic-specific script features like contextual letter forms, ligatures and dots, and optional diacritics. The paper will suggest and mathematically model a confidence-aware hybrid retrieval system in Arabic document images with the use of keyword queries. The framework integrates a lexical branch (OCR-based), a visual region-matching branch, branch-based score calibration, confidence gated fusion, re-ranking and keyword localization. The empirical part is limited by the measures which are directly justified by the measured datasets. On the FUNSD testing split (50 pages; 668 normalized single-token searches chosen in ground-truth annotations) an experiment of keyword retrieval at the page level was implemented. Using Tesseract 5.5.0 English OCR and BM25 indexing, OCR-BM25 achieved P@1 = 0.867, mAP = 0.678, NDCG@10 = 0.713, and MRR = 0.888; the confidence-weighted OCR variant achieved P@1 = 0.870, mAP = 0.672, NDCG@10 = 0.708, and MRR = 0.890. The given measured values confirm the auxiliary lexical retrieval protocol of non-Arabic noisy forms and demonstrate that raw OCR confidence is not uniformly beneficial when it comes to ranking metrics. They are not presented as Arabic hybrid retrieval performance. The assessed IFN/ENIT files include Arabic handwritten word images and segmentation XML but do not include page-level keyword relevance labels, whereas assessed RVL-CDIP test folders include document-class images instead of keyword annotations. It is based on this that the paper presents Arabic-oriented framework, reproducible evaluation protocol, measured auxiliary evidence and the identification of the dedicated Arabic retrieval benchmark that will be needed to complete target-domain validation.

Read PDF