PURPOSE
This study aimed to evaluate how speech signal compression algorithms affect the acoustic and perceptual characteristics of dysarthric speech. As telepractice becomes more common in speech-language pathology, particularly in the assessment of dysarthria, understanding the impact of telecommunication compression on speech fidelity is essential for clinical decision making.
METHOD
Speech samples were collected from 15 individuals with dysarthria, recorded locally and simultaneously via Zoom (Version 6.2.11). The Opus codec was used to process the locally recorded speech samples under fullband, wideband, and narrowband compression conditions. Acoustic measures were derived for all samples. Twenty experienced speech-language pathologists (SLPs) rated vocal quality and nasality for sustained vowels and rated articulatory precision and orthographically transcribed sentences across each condition.
RESULTS
Narrowband compression was associated with significantly degraded voice quality measures (e.g., harmonic-to-noise ratio, shimmer) and reduced transcription accuracy. Despite these changes in acoustics and intelligibility, ratings of articulatory precision, nasality, and vocal quality remained stable across conditions, suggesting SLPs could adjust for compression artifacts when making perceptual judgments.
CONCLUSIONS
Speech compression, particularly in low-bandwidth (i.e., narrow) condition, impacts acoustic fidelity and intelligibility, though expert listeners maintain reliable perceptual judgments. These findings underscore the need to consider compression effects in telepractice and highlight the importance of developing optimized protocols for remote dysarthria assessment.
Kelvin Tran, Rene L. Utianski, V. Berisha et al.· Journal of Speech, Language...· 0 citations
Consonants contribute unequally to whether a word is understood. Given the limited time available for therapy, ranking consonants by contribution to intelligibility helps prioritize intervention targets in motor speech disorders. However, measuring this contribution relies on perceptual studies that are difficult to scale. This paper presents a scalable method that measures consonant contribution using acoustic masking. We silence one consonant at a time in an isolated word and test whether an automatic speech recognition (ASR) model still recognizes the word. We define a consonant's contribution score as the proportion of its masked instances for which the word becomes misrecognized, which we refer to as the mask-induced misrecognition rate (MMR). We relate MMR to two linguistic factors previously reported to correlate with consonant contribution, namely phoneme frequency and functional load. We apply this analysis across four languages, English, Spanish, German, and Czech, using three ASR architectures, MMS (encoder-only), Whisper (encoder-decoder), and Qwen3-ASR (LLM-based). Using partial Spearman correlations, we find that phoneme frequency correlates negatively with MMR while functional load correlates positively. In other words, more frequent consonants are less disruptive when masked, whereas consonants carrying more lexical contrast are more disruptive. Further cross-language analysis shows that consonant rankings agree only partially across languages, indicating that consonant contribution is language-dependent.
Eunjung Yeo, Kwanghee Choi, K. Kothadia et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.