Skip to content
Open access

Towards the Automated Recognition of Ancient Sundanese Scripts: Groundtruth Generation at the Geusan Ulun Museum

Aug 2026 · Engineering, Technology & Applied Science Research · 0 citations · 13 references

Abstract

While cultural preservation and the digitization of historical manuscripts have a long history, low-resource writing systems continue to be poorly represented due to the absence of reliable ground truth datasets. This paper describes a reproducible engineering process for building a ground truth-validated dataset for deteriorated Old Sundanese manuscripts from the Geusan Ulun Museum in Indonesia. The Old Sundanese Manuscripts Dataset (OSMD) was constructed by combining controlled digitization, contrast enhancement, edge-based object extraction, morphological image processing, expert annotation, and optimized dataset packaging. The dataset contains 1,300 word-level samples for five consonant classes: Ha, Na, Ca, Ra, and Ka. To assess the potential of the dataset, baseline classification experiments were conducted using HOG+SVM and a lightweight CNN, where the latter outperformed (p<0.01) with 0.842 accuracy and 0.833 macro F1-score. The absence of contrast enhancement resulted in a notable drop in performance, and cross-validation demonstrated low variance (±0.006). The process followed in this study can be applied to low-resource historical scripts, connecting cultural heritage and computational modeling.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.