This work proposes an acoustics-based assessment approach that employs a convolutional neural network to extract rhythm features directly from the speech amplitude envelope, motivated by evidence linking low-frequency modulations to rhythm perception.
Abstract
Automated Speaking Assessment of non-native speech must effectively evaluate prosody, including speech rhythm, to align with human perception. However, commonly employed rhythm metrics rely on segmental duration, requiring an additional alignment step, which is error-prone in non-native speech containing disfluencies and mispronunciations. We propose an acoustics-based assessment approach that employs a convolutional neural network to extract rhythm features directly from the speech amplitude envelope, motivated by evidence linking low-frequency modulations to rhythm perception. The proposed models are trained on a proficiency score regression task using the speechocean762 dataset and compared against duration-based models. Our results show that a model using the amplitude envelope's first derivative achieves the highest correlation with human-assigned scores on the Fluency and Prosody dimensions, producing significantly lower errors than one using segment durations among less fluent speakers. The findings support acoustic envelope features as robust, alignment-free alternatives for L2 rhythm assessment. Code is released publicly.
CASA, a simpler architecture combining Whisper-medium and Qwen3.5-2B that achieves state-of-the-art performance while providing a more interpretable separation between speech delivery and content, is proposed.
Nhan Phan, Ilona Lähteenmäki, Anna von Zansen et al.· 0 citations
An interpretable audio-visual synchronization framework that explicitly models speech-lip temporal alignment through feature encoding, cross-attention fusion, lag estimation, stability quantification, and visualization is proposed.
Bo-Wen Yu, Ming-Yue Huang, Yi-Shen Liu et al.· 0 citations
Background In background noise, speakers adapt their speech production, giving rise to Lombard speech, which often improves speech intelligibility (SI). While intelligibility benefits of Lombard speech have been extensively studied in non-tonal languages, it remains unclear whether spectro-temporal cues, which are crit...
M. Scharf, A. Warzybok, L. Wong et al.· PLoS ONE· 0 citations
This work demonstrates TAPA on the 2016 U.S. presidential debate, and suggests that TAPA can be used to increase access to naturalistic speech data and speed up the processing timeline with experts' supervision.
Ethan Kutlu, Emerson Peters, Ciara Tapanes et al.· Behavior Research Methods· 0 citations
Speech assessment tools for speakers with speech pathology must be both accurate and interpretable if they are to be adopted in clinical practice. Existing reference-audio measures such as the Neural Acoustic Distance (NAD) reach high speaker-level correlations with listener intelligibility scores but operate on self-s...
B. Halpern, Thomas B. Tienkamp, D. Abur et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.