Skip to content
Preprint

Automated Assessment of L2 Speech Rhythm Using Low-Frequency Amplitude Modulations

Sep 2026 · 0 citations · 27 references
Computer Science

TL;DR

This work proposes an acoustics-based assessment approach that employs a convolutional neural network to extract rhythm features directly from the speech amplitude envelope, motivated by evidence linking low-frequency modulations to rhythm perception.

Abstract

Automated Speaking Assessment of non-native speech must effectively evaluate prosody, including speech rhythm, to align with human perception. However, commonly employed rhythm metrics rely on segmental duration, requiring an additional alignment step, which is error-prone in non-native speech containing disfluencies and mispronunciations. We propose an acoustics-based assessment approach that employs a convolutional neural network to extract rhythm features directly from the speech amplitude envelope, motivated by evidence linking low-frequency modulations to rhythm perception. The proposed models are trained on a proficiency score regression task using the speechocean762 dataset and compared against duration-based models. Our results show that a model using the amplitude envelope's first derivative achieves the highest correlation with human-assigned scores on the Fluency and Prosody dimensions, producing significantly lower errors than one using segment durations among less fluent speakers. The findings support acoustic envelope features as robust, alignment-free alternatives for L2 rhythm assessment. Code is released publicly.

View source

Similar papers

Preprint Aug 2026

CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model

CASA, a simpler architecture combining Whisper-medium and Qwen3.5-2B that achieves state-of-the-art performance while providing a more interpretable separation between speech delivery and content, is proposed.

Nhan Phan, Ilona Lähteenmäki, Anna von Zansen et al. · 0 citations
Open access Aug 2026

Spectro-temporal vs. spectral features to predict the lombard gain in Mandarin Chinese

Background In background noise, speakers adapt their speech production, giving rise to Lombard speech, which often improves speech intelligibility (SI). While intelligibility benefits of Lombard speech have been extensively studied in non-tonal languages, it remains unclear whether spectro-temporal cues, which are crit...

M. Scharf, A. Warzybok, L. Wong et al. · 0 citations
Open access Sep 2026

Toolkit for acoustic–phonetic analysis of naturalistic speech data

This work demonstrates TAPA on the 2016 U.S. presidential debate, and suggests that TAPA can be used to increase access to naturalistic speech data and speed up the processing timeline with experts' supervision.

Ethan Kutlu, Emerson Peters, Ciara Tapanes et al. · 0 citations
Preprint Sep 2026

ART-NAD: An Articulatory Inversion-based Neural Acoustic Distance for Pathological Speech Intelligibility Assessment

Speech assessment tools for speakers with speech pathology must be both accurate and interpretable if they are to be adopted in clinical practice. Existing reference-audio measures such as the Neural Acoustic Distance (NAD) reach high speaker-level correlations with listener intelligibility scores but operate on self-s...

B. Halpern, Thomas B. Tienkamp, D. Abur et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.