A Cascade English–Vietnamese Speech-to-Speech Translation Framework for Online IT Education
Abstract
The rapid expansion of online IT education has created a pressing need to bridge the English-Vietnamese language barrier for non-native learners, as subtitle-based solutions impose cognitive load and fail to preserve the lecturer’s vocal identity. This paper presents a cascade Speech-to-Speech Translation (S2ST) pipeline for IT lectures, combining Automatic Speech Recognition (ASR), an LLM-based Machine Translation (MT) backend selected against conventional NMT baselines, and zero-shot Text-to-Speech (TTS) with voice cloning. We construct a domain-specific dataset of 240 English IT lecture videos and a 40-video human-annotated benchmark. Across five ASR models, five MT models, and two TTS models, Whisper Medium achieves the lowest Word Error Rate (3.36%), Gemini 2.0 Flash the highest translation quality (BLEU 55, chrF 72), and F5-TTS the most faithful voice cloning (Speaker Similarity 0.782, MOS 4.0). In an indicative single-lecture comparison, the cascade pipeline reaches an MOS of 4.2 versus 2.2 for the SeamlessM4T-Large end-to-end baseline. An in-depth error analysis and practical mitigation strategies further suggest that a carefully engineered cascade remains a competitive choice for low-resource educational S2ST.