Skip to content
Conference

Benchmarking Open-Source Vietnamese-English Speech-to-Text Translation Systems

Aug 2026 · International Conference on Multimedia Analysis and Pattern Recognition · pp. 718-723 · 0 citations · 28 references

Abstract

Speech-to-text translation for low-resource language pairs such as Vietnamese–English remains underexplored, despite growing demand in real-world applications. In this study, we present a systematic zero-shot benchmark on the Vietnamese test split of FLEURS, evaluating 30 configurations across three system families: commercial APIs, open-source cascade pipelines, and open-source end-to-end models. Our cascade systems pair four ASR backbones—multilingual Whisper and Vietnamese-fine-tuned PhoWhisper variants—with six MT backends spanning dedicated translation models and instruction-tuned LLMs. We find that ASR quality places a clear upper bound on downstream translation, while MT backend selection has at least a comparable effect on output quality. The strongest open-source cascade, Whisper-large + NLLB-3.3B, reaches 30.03 BLEU, approaching the commercial Whisper-1 + GPT-5.4 reference at 31.11 BLEU, whereas open-source end-to-end models lag substantially behind in translation quality. A joint quality–efficiency analysis shows that open-source end-to-end models are about 2–4× faster than most open-source cascades, while LLM-based MT backends incur considerable latency overhead for modest quality gains. Several LLM-based backends, especially Gemma-2-9B, produce semantically faithful but lexically divergent translations relative to specialized MT models, as evidenced by higher COMET but lower BLEU scores. These results suggest that open-source components are approaching commercial cascade quality on this benchmark, while closing the end-to-end gap remains the key challenge for open-source Vi→En speech translation.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.