A data-centric approach to performance improvement in under-resourced ASR: The case of Dënë Sųłıné
Abstract
This paper presents a study focused on advancing Automatic Speech Recognition (ASR) for the under-resourced language Dënë Su ˛ łıné through data-centric approaches. We explore multiple strategies to enhance the quality of training data—both audio recordings and tran-scriptions—to address the challenges posed by mixed-quality datasets. Our experiments investigate which data preparation techniques most effectively improve ASR performance in this context. Our findings show that reducing spelling variants of the same lexeme in the corpus significantly improves model generalization, resulting in a substantial increase in recognition accuracy. Additionally, we demonstrate that increasing manually reviewed transcriptions consistently improves word and character error rates, while audio enhancement slightly reduces performance, highlighting the complex trade-offs in low-resource ASR development.