Skip to content
Conference

A Cascade English–Vietnamese Speech-to-Speech Translation Framework for Online IT Education

Aug 2026 · International Conference on Multimedia Analysis and Pattern Recognition · pp. 664-669 · 0 citations · 26 references

Abstract

The rapid expansion of online IT education has created a pressing need to bridge the English-Vietnamese language barrier for non-native learners, as subtitle-based solutions impose cognitive load and fail to preserve the lecturer’s vocal identity. This paper presents a cascade Speech-to-Speech Translation (S2ST) pipeline for IT lectures, combining Automatic Speech Recognition (ASR), an LLM-based Machine Translation (MT) backend selected against conventional NMT baselines, and zero-shot Text-to-Speech (TTS) with voice cloning. We construct a domain-specific dataset of 240 English IT lecture videos and a 40-video human-annotated benchmark. Across five ASR models, five MT models, and two TTS models, Whisper Medium achieves the lowest Word Error Rate (3.36%), Gemini 2.0 Flash the highest translation quality (BLEU 55, chrF 72), and F5-TTS the most faithful voice cloning (Speaker Similarity 0.782, MOS 4.0). In an indicative single-lecture comparison, the cascade pipeline reaches an MOS of 4.2 versus 2.2 for the SeamlessM4T-Large end-to-end baseline. An in-depth error analysis and practical mitigation strategies further suggest that a carefully engineered cascade remains a competitive choice for low-resource educational S2ST.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.