Towards the Automated Recognition of Ancient Sundanese Scripts: Groundtruth Generation at the Geusan Ulun Museum
Abstract
While cultural preservation and the digitization of historical manuscripts have a long history, low-resource writing systems continue to be poorly represented due to the absence of reliable ground truth datasets. This paper describes a reproducible engineering process for building a ground truth-validated dataset for deteriorated Old Sundanese manuscripts from the Geusan Ulun Museum in Indonesia. The Old Sundanese Manuscripts Dataset (OSMD) was constructed by combining controlled digitization, contrast enhancement, edge-based object extraction, morphological image processing, expert annotation, and optimized dataset packaging. The dataset contains 1,300 word-level samples for five consonant classes: Ha, Na, Ca, Ra, and Ka. To assess the potential of the dataset, baseline classification experiments were conducted using HOG+SVM and a lightweight CNN, where the latter outperformed (p<0.01) with 0.842 accuracy and 0.833 macro F1-score. The absence of contrast enhancement resulted in a notable drop in performance, and cross-validation demonstrated low variance (±0.006). The process followed in this study can be applied to low-resource historical scripts, connecting cultural heritage and computational modeling.