Non-Parallel Dysarthric Speech Enhancement Using CycleGAN-VC2 and Melgan with Ablation Validation
Abstract
Deploying dysarthric speech enhancement in clinical settings runs into a hard practical wall: collecting matched recordings from patients across multiple sessions is rarely achievable because medication timing, motor fatigue, and dayto-day articulatory variability work against the frame-aligned parallel corpora that conventional voice conversion (VC) systems require. This paper presents a pipeline designed specifically to bypass that constraint, learning entirely from unpaired recordings with no utterance matching at any stage. The system chains CycleGAN-VC2 spectral conversion with MelGAN waveform synthesis, followed by ITU-R BS. 1770 integrated loudness normalisation a post-processing step absent from most comparable pipelines that removes a measurement confound capable of inflating published objective scores while also ensuring consistent output level for assistive device deployment. Evaluated on the UASpeech benchmark using CER, WER, STOI, MCD, and MOS, the proposed system achieves a 43% relative CER reduction (31.4%→17.9%), improves STOI from 0.66 to 0.84, reduces MCD from 7.8 dB to 5.1 dB, and raises MOS from 2.1 to 3.8. All improvements are statistically significant ($p<0.01$, paired Wilcoxon). A three-condition ablation with statistical testing shows that CycleGAN-VC2 primarily drives intelligibility gain while MelGAN drives perceived naturalness, and that the two contributions are largely orthogonal explaining why the full system leads every single-component and baseline configuration on every evaluation axis simultaneously.