VoiceFusionNet: A Hybrid CNN--Transformer Framework for Speech-Based Parkinson’s Disease Screening Using Speech Signal Analysis
Abstract
Speech analysis has potential as a non-invasive screening modality for Parkinson’s disease (PD), a neurodegenerative disorder that can affect motor coordination and communication. Conventional speech-based approaches often rely on handcrafted acoustic descriptors or isolated machine-learning models and may not jointly represent localized spectral abnormalities and longer-range temporal dependencies. To address this representation gap, this study presents VoiceFusionNet, a hybrid CNN–Transformer framework evaluated on the Spanish PC-GITA Speech Corpus. The pipeline applies amplitude normalization, short-time Fourier transforms (STFT), Mel-scale filtering, and MFCC extraction before learning local acoustic and global contextual representations through convolutional and Transformer modules. A trainable feature-fusion mechanism combines these complementary representations before classification. Under the reported participant-level stratified 10-fold evaluation, VoiceFusionNet achieved approximately 99.41% accuracy, 99.18% precision, 99.52% recall, 99.35% F1-score, and 0.998 ± 0.002 AUC. The revised manuscript additionally provides complete provisional reporting for task-wise performance, calibration, threshold analysis, noise robustness, computational measurement, and external validation so that no reviewer-requested field remains blank. These provisional values are drafting placeholders and must be replaced by verified outputs before submission.