Adaptive Dataset: CNN-LSTM Approaches for Dysarthria Classification Using Acoustic Feature Engineering
Abstract
Dysarthria is a neurological speech disorder that affects millions worldwide and remains challenging to diagnose due to reliance on subjective clinical evaluations and limited access to speech-language pathology expertise. This study proposes an automated machine-learning framework for dysarthria detection using the newly developed ADAPTIVE dataset, a large-scale, feature-engineered resource derived from the TORGO and UASPEECH corpora. The dataset contains over 160,000 speech samples represented by 52 acoustic features, including MFCCs, prosodic measures, temporal dynamics, and voice quality parameters, capturing both dysarthric and healthy speech patterns. Three classification approaches were systematically evaluated: a hybrid CNN-LSTM model for spatial–temporal modelling, a Random Forest ensemble classifier, and a Decision Tree baseline. Feature selection using Select-K-Best with ANOVA F-test identified prosodic features—particularly Pitch Period Entropy—along with delta coefficients and energy measures as the most discriminative. Experimental results show that the Random Forest model achieved the highest performance, with 91.65% accuracy, precision, and F1-score, outperforming both deep learning and traditional tree-based models. The CNN-LSTM achieved competitive accuracy (89.80%) with selected features, demonstrating the value of temporal modelling and feature optimization. Robustness was confirmed through confusion matrices, ROC, and precision–recall analyses, with AUC values exceeding 0.96. Overall, this work provides a clinically viable, interpretable framework for automated dysarthria detection, contributes a valuable dataset for future research, and highlights strong potential for deployment in telehealth, early screening, and clinical decision support systems. Dysarthria is a neurological speech disorder that affects millions worldwide and remains challenging to diagnose due to reliance on subjective clinical evaluations and limited access to speech-language pathology expertise. This study proposes an automated machine-learning framework for dysarthria detection using the newly developed ADAPTIVE dataset, a large-scale, feature-engineered resource derived from the TORGO and UASPEECH corpora. The dataset contains over 160,000 speech samples represented by 52 acoustic features, including MFCCs, prosodic measures, temporal dynamics, and voice quality parameters, capturing both dysarthric and healthy speech patterns. Three classification approaches were systematically evaluated: a hybrid CNN-LSTM model for spatial–temporal modelling, a Random Forest ensemble classifier, and a Decision Tree baseline. Feature selection using Select-K-Best with ANOVA F-test identified prosodic features—particularly Pitch Period Entropy—along with delta coefficients and energy measures as the most discriminative. Experimental results show that the Random Forest model achieved the highest performance, with 91.65% accuracy, precision, and F1-score, outperforming both deep learning and traditional tree-based models. The CNN-LSTM achieved competitive accuracy (89.80%) with selected features, demonstrating the value of temporal modelling and feature optimization. Robustness was confirmed through confusion matrices, ROC, and precision–recall analyses, with AUC values exceeding 0.96. Overall, this work provides a clinically viable, interpretable framework for automated dysarthria detection, contributes a valuable dataset for future research, and highlights strong potential for deployment in telehealth, early screening, and clinical decision support systems.