Skip to content
Open access

Automatic Detection of Misarticulation in Low-Resource Language Children Using Kaldi-Based ASR and Machine Learning Approaches

2026 · IEEE Access · Vol 14, pp. 140716-140728 · 0 citations · 45 references

Abstract

The lack of specific standards for evaluating the significant variation in children’s speech and the limited availability of annotated speech samples make it hard to automatically identify misarticulation in children speaking low-resource Indian languages. While techniques for assessing pronunciation using Automatic Speech Recognition (ASR) have shown positive results for high-resource languages, there has been little focus on applying them to identify phoneme-level misarticulation in Marathi children’s speech. Marathi is considered one of the major Indo-Aryan languages spoken by millions of people in India. But in speech and language technology research, it remains underrepresented as compared to high-resource languages such as English. This research proposes a new hybrid framework that combines machine learning methods with Kaldi-based ASR to automatically detect misarticulated speech in children who speak Marathi. The study used speech samples from children aged five to fifteen, including both typically developing speakers and those with articulation challenges, to create a specialized speech corpus. The Kaldi toolbox was employed to develop a Marathi ASR system with acoustic and linguistic models tailored for phoneme-level analysis. The proposed method combines forced alignment, phoneme location information, Goodness of Pronunciation (GOP) ratings, MFCC statistical characteristics, and phonetic confidence metrics based on posterior probability to capture articulation error more precisely. The classification of phonemes as correctly or incorrectly articulated using these qualities is trained and assessed using various machine learning methods, including K-Nearest Neighbors (KNN), Support Vector Machines (SVM), and ensemble models. Empirical results confirmed that combining ASR-derived phonetic features with data-driven machine learning techniques is effective for resource-constrained speech therapy applications. For the early detection of speech sound issues in Marathi children, the suggested architecture provides an unbiased, scalable, and adaptable method. This work enhances speech assessment techniques for under-resourced Indian languages and opens the door for future AI-supported speech therapy and clinical evaluation tools.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.