Speech-Based Alzheimer's Detection Under Controlled Acoustic Perturbations: A Baseline Study of Class Imbalance and Robustness
Abstract
Background: Alzheimer's disease (AD) affects over 55 million people worldwide, and early detection through speech analysis offers a low-cost, non-invasive alternative to neuroimaging. However, clinical deployment of speech-based classifiers faces underexplored challenges: domain shifts across recording environments, limited labeled data at new sites, and a need for calibrated uncertainty estimates. Methods: We evaluated Random Forest and XGBoost classifiers on DementiaBank (295 Cookie Theft recordings from 292 unique speakers; 45 control, 250 AD). We extracted 1,589 features combining 53 hand-crafted acoustic descriptors (librosa, Praat) with wav2vec 2.0 and HuBERT embeddings, reduced to 128 dimensions via PCA. Robustness was assessed using 8 controlled acoustic simulations (microphone variation, background noise at 10-25 dB SNR, and reverberation at RT60 0.3-0.8 s) calibrated to published clinical measurements. Results: Both models achieved 81.4% raw accuracy (48/59; macro F1: 0.594; balanced accuracy: 0.581) on clean data, below the majority-class raw baseline of 84.7% but above it in balanced accuracy (0.500) and macro F1 (0.459). Under 7 of 8 controlled acoustic simulations, raw accuracy converged to 84.7%, matching the majority-class rate. A supplementary confusion-matrix analysis using a simplified verification model confirmed that this pattern reflects majority-class defaulting: control recall dropped to 0.000 under microphone and noise simulations. Random Forest achieved low calibration error (ECE = 0.042) on the 59-sample test set. In a single-draw few-shot experiment, XGBoost yielded 67.8% with K=10 examples per class. Conclusion: On clean data, acoustic features combined with self-supervised embeddings produce classifiers with meaningful balanced performance despite severe class imbalance (85% AD). The controlled acoustic simulation framework provides a reproducible evaluation methodology, though our results reveal that apparent robustness may reflect majority-class defaulting rather than genuine invariance, an important methodological caution for future work.