Skip to content
Conference

Benchmarking Speech Translation for Hindi and Tamil Using Pretrained Models

Aug 2026 · International Conference on Information Security and Cryptology · pp. 413-419 · 0 citations · 21 references

Abstract

Artificial Intelligence has significantly boosted the process of Machine Translation due to its rapid advancement. In this paper we present a study on the ability of pre-trained deep learning models for speech translation when applied in intelligent and industrial communication systems. There are significant advances in automatic speech recognition (ASR) and speech-totext systems at this time. The extensive use of data for training such models has been proved to be the key to the most recent advances in speech transcription. However, some of these tools have not achieved the same degree of accuracy when it comes to some Indian languages. While, for example, Whisper or Wav2Vec 2.0 achieve good performance on speech-to-text conversion and multilingual efforts such as NLLB are making strides beyond language boundaries; issues remain, particularly for less commonly spoken indigenous languages. With the advent of speech to speech translation, evaulation of its performance continues to be challenging- due to the reliance upon reference translations, which are scarce for many Indian languages. This is an investigation to see if Whisper performs well with Hindi and Tamil speech and Wav2Vec 2.0. Different segments (in length and flow) from LibriSpeech were used as audio input for both systems. Score is computed by word error rate and BLEU, and human review having closer look at the translation quality. While numbers may help quantify output, meaning is more likely to be evident when a person reads the results. Where Wav2Vec 2.0 fails to sustain, Whisper holds up: particularly for sounds that linger, that is. One falters with changes in pitch or changes in loudness; the other can keep up. Garbage in is garbage out - errors early in the process of converting speech to text can only get worse in the process of converting the latter to speech. The ability to process everything simultaneously saves times, which can be crucial when timing is a critical. Pre-built AI systems have demonstrated their ability to transcend languages, seamlessly integrating into active workspaces where conversations can freely flow.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.