Acoustic-based speech chunking techniques for mild cognitive impairment classification
Abstract
Early and timely intervention is critical for the detection of mild cognitive impairment (MCI) to reduce the progression of dementia. However, traditional diagnostic methods remain time-consuming, resource-intensive, and difficult to scale. Speech-based voice analysis offers a non-invasive and accessible alternative, providing insights into an individual’s cognitive processes and enabling data collection in real-world settings. While recent advances in deep learning have improved acoustic representation learning, the role of preprocessing, particularly speech chunking, remains underexplored. In this thesis, we investigate the impact of speech chunking and aggregation techniques on MCI detection using learned acoustic representations. We extract 1280-dimensional acoustic embeddings using the Whisper-Large-v3 encoder from multilingual TAUKADIAL speech recordings by applying speech chunking with three chunk sizes {5s, 15s, 30s}, four overlap ratios {0, 0.25, 0.5, 0.75}, and two chunk aggregation methods (mean and max pooling). We evaluate the framework using two MLP models and two prediction aggregation methods (majority voting and probability averaging) across 10 training strategies under unilingual, multilingual, and cross-lingual regime. The results show that chunk size is a critical factor in representation quality, with 15-second chunks providing the best balance between stability and performance. Performance degrades as the overlap ratio increases, while mean pooling yields more stable and consistent results than max pooling. Additionally, simpler models outperform deeper architectures, with negligible differences observed between prediction aggregation methods. Language alignment emerges as the primary performance bottleneck: unilingual training achieves the best results, cross-lingual generalization remains weak, and multilingual training improves robustness. In summary, this thesis demonstrates that speech chunking is not merely a preprocessing step but a key component of representation design, offering practical insights for developing robust, scalable, and language-aware speech-based systems for early detection of cognitive impairment.