Non-Invasive Parkinson's Disease Screening from Vocal Biomarkers: A Subject-Level Machine Learning Framework with Bayesian Optimisation and Explainable AI
Abstract
An estimated ten million people live with Parkinson's, yet timely diagnosis remains constrained by specialist assessment, costly imaging, and unequal care access. Phonation analysis offers a low-cost alternative: dopaminergic degeneration causes vocal impairments years before motor symptoms emerge, enabling community screening. However, existing ML approaches are limited by recording-level data leakage and opacity. This paper presents a four-phase acoustic framework-spanning signal acquisition, feature processing, and interpretable decision support-applied to a multi-type Parkinson's speech dataset (40 training, 28 blind-test). The framework enforces subject-level partitioning via GroupKFold to eliminate leakage, derives 104 participant-level acoustic features through multi-statistic aggregation (mean, median, standard deviation, interquartile range), and employs Bayesian optimisation using Optuna. The optimised SVM achieved 90.00% cross-validated accuracy, a 12.50 percentage-point improvement over the 77.50% benchmark. Cross-validated sensitivity was 95.00%, specificity 85.00%, Youden Index \(J=0.80\), Brier Score \(\mathrm{BS}=0.1049\). External validation on a held-out vowel-only cohort achieved 85.71% sensitivity, surpassing the benchmark. A permutation test (\(p=0.349\)) indicated no significant linear mapping between acoustic features and motor severity, motivating binary classification. Shapley Additive Explanations identified interquartile range of degree-of-voice-breaks and shimmer amplitude variability as primary diagnostic drivers, aligning with neuroacoustic evidence for aperiodic phonation. These results show that leakage prevention, variance-sensitive engineering, and transparent attribution together constitute a viable foundation for remote, telehealth-linked pre-screening instruments in resource-constrained settings.