Skip to content
Preprint

When Speech Meets Lips: Interpretable Audio-Visual Synchronization for L2 Pronunciation Assessment

Sep 2026 · 0 citations · 42 references
Computer Science

TL;DR

An interpretable audio-visual synchronization framework that explicitly models speech-lip temporal alignment through feature encoding, cross-attention fusion, lag estimation, stability quantification, and visualization is proposed.

Abstract

Automatic Pronunciation Assessment (APA) systems have achieved strong performance with transformer-based models and self-supervised speech representations. However, most methods rely only on acoustic signals and overlook temporal synchronization between speech and articulatory movements, limiting diagnostic feedback on timing mismatches important for L2 pronunciation training. We propose an interpretable audio-visual synchronization framework that explicitly models speech-lip temporal alignment through feature encoding, cross-attention fusion, lag estimation, stability quantification, and visualization. The framework introduces frame-level lag trajectories and a Lag Stability Index (LSI) to quantify synchronization robustness. We also interviewed 30 participants, including 10 instructors and 20 students with diverse first-language backgrounds, to assess its effectiveness. By transforming implicit alignment into interpretable representations, the framework connects automatic scoring with actionable Computer-Aided Pronunciation Training feedback. Datasets and supplemental materials are available at https://www.robots.ox.ac.uk/~vgg/data/lip_reading/.

View source

Similar papers

Preprint Sep 2026

Automated Assessment of L2 Speech Rhythm Using Low-Frequency Amplitude Modulations

This work proposes an acoustics-based assessment approach that employs a convolutional neural network to extract rhythm features directly from the speech amplitude envelope, motivated by evidence linking low-frequency modulations to rhythm perception.

João Lima, Lucas H. Ueda, P. Costa · 0 citations
Open access Aug 2026

Intelligent Interpretation of Traditional Folk Tales through Vocalization: a Generative Model Based on Emotion-Driven and Lip-Syncing

This study presents an integrated multi-modal signal generation framework for intelligent vocalization of traditional folk tales, ensuring emotion-driven speech synthesis with precise lip synchronization. The proposed system models continuous narrative emotion trajectories and embeds them into phoneme-level acoustic ge...

S. Li · 0 citations
Preprint Sep 2026

Subphonetic Acoustic Modeling via Optimal Transport for Pronunciation Assessment

Pronunciation assessment requires acoustic evidence that is temporally precise, diagnostically meaningful, and faithful to the learner's actual production. However, existing acoustic models often struggle to provide recognition and segmentation evidence simultaneously. CTC-based phone recognizers can predict phone sequ...

Hao-Peng Geng, Jiun-Ting Li, Daisuke Saito et al. · 0 citations
Preprint Sep 2026

ART-NAD: An Articulatory Inversion-based Neural Acoustic Distance for Pathological Speech Intelligibility Assessment

Speech assessment tools for speakers with speech pathology must be both accurate and interpretable if they are to be adopted in clinical practice. Existing reference-audio measures such as the Neural Acoustic Distance (NAD) reach high speaker-level correlations with listener intelligibility scores but operate on self-s...

B. Halpern, Thomas B. Tienkamp, D. Abur et al. · 0 citations
Open access Sep 2026

Toolkit for acoustic–phonetic analysis of naturalistic speech data

This work demonstrates TAPA on the 2016 U.S. presidential debate, and suggests that TAPA can be used to increase access to naturalistic speech data and speed up the processing timeline with experts' supervision.

Ethan Kutlu, Emerson Peters, Ciara Tapanes et al. · 0 citations
Preprint Sep 2026

Align Then Reason: A Multimodal Lip-Sync Judge for Dubbing

Dubbing quality control requires a reference-free judge that can determine whether a candidate text line matches a speaker's visible articulation in both content and timing, using only silent video and text because dubbed audio may not yet exist. Existing visual speech recognizers and video-language models are poorly s...

Rui Liu, Bhavin Jawade, Hao-Qi Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.