Skip to content
Preprint

A Speech Corpus for Mizo Automatic Speech Recognition: Whisper and SraVaani 1.0 Fine-Tuning with Morphology-Aware Evaluation

Aug 2026 · 0 citations · 9 references
Computer Science Engineering

TL;DR

The results demonstrate that the Whisper model can achieve a substantially low WER, even when adapted to an unseen language.

Abstract

This study reports the development of an Automatic Speech Recognition (ASR) system in Mizo, a low-resource language. The development included collecting 17.62 hours of speech data, curating it, and fine-tuning the Mizo ASR system with three Whisper multilingual models and with the SraVaani 1.0 Indic multilingual model. Whisper-large-v3 achieved the lowest conventional WER (18.08%), while morphology-aware evaluation yielded a WER of 7.22%. Zero-shot evaluation of the SraVaani 1.0 Indic multilingual model yielded a WER of 58.27%, while Mizo-specific fine-tuning reduced the conventional WER to 29.45% and the morphology-aware WER to 17.93%. The results demonstrate that the Whisper model can achieve a substantially low WER, even when adapted to an unseen language. In contrast, SraVaani 1.0 supports the Mizo language in its multilingual model; however, fine-tuning with carefully curated Mizo speech data substantially improves its performance.

View source

Similar papers

#small language model Preprint Aug 2026

Fine-Tuning Whisper for Automatic Speech Recognition in Baniwa: A Preliminary Study

A preliminary study on the adaptation of Whisper for Automatic Speech Recognition in Baniwa, an indigenous Arawakan language spoken in Brazil, Colombia, and Venezuela, demonstrating that multilingual foundation models can be successfully adapted to extremely low-resource indigenous languages.

Leonardo Duart, T. Fonseca, T. Chacon · 0 citations
Open access Aug 2026

Speech-to-text model comparison using XLS-R, XLSR-53, and Wav2Vec 2.0

Findings indicate that language-specific fine-tuning plays a more critical role than multilingual generalization in achieving accurate ASR for Indonesian and provide practical guidance for deploying ASR systems in low-resource language scenarios.

J. Hebert, Amalia Zahra · 0 citations
Preprint Aug 2026

SraVaani 1.0: Scaling Inclusive Speech Recognition for Indic Languages

SraVaani-1.0 achieves the lowest word error rate (WER) on a large number of language-dataset pairs while remaining competitive with the best-performing systems on high resource while being assessed exclusively on the VAANI benchmark.

Sujith Pulikodan, A. Basu, J. Pavankumar et al. · 1 citation · ⚡1
Open access Jul 2026

A comparative analysis of pretrained Wav2Vec XLSR-53 and Whisper-Small models for automatic speech recognition in the Telugu language

This paper investigates the effectiveness of pre-trained models like Wav2Vec XLSR-53 and Whisper-Small for developing ASR systems for the Telugu language, addressing the challenge of limited data availability and demonstrating satisfactory results even when fine-tuned on a smaller dataset.

J. Pushparaj, Muzaffar Ahmad Dar, Sri Gani Kaarthikeya Kammula et al. · 0 citations
#natural language process... Preprint Sep 2026

BuzzASR: A Swarm of 100+ Monolingual Speech Recognition Models

We introduce BuzzASR, a collection of language-specialized fine-tuned Whisper models adapted for automatic speech recognition (ASR) in 102 languages. Large end-to-end Transformer-based ASR models such as Whisper have revolutionized ASR, but most prominent models are highly multilingual. As a result, these models often perform poorly on languages less well-represented in their training set. While it has long been known that effective language adaptation can be achieved through simple fine-tuning on monolingual data, this strategy has only been applied to a small number of languages. We massively scale up this simple approach to 102 languages covered in the FLEURS dataset, while also implementing a more complex language adaptation strategy that integrates monolingual tokenizer replacement and data augmentation using text-only fine-tuning. BuzzASR models outperform Whisper-large-v3 on 77 out of 102 languages, reducing character error rates (CER) by a factor of over 2.8 on average. Our models achieve state-of-the-art CER among open-source systems on 27 of 102 languages on the combined FLEURS and Common Voice test set. Our tokenizer replacement strategy yields an average 3.3x improvement in compression rate (characters per token) over Whisper's multilingual BPE, with gains of up to 21.7x. We release all models, code, and detailed results: https://lemn-lab.github.io/buzz-asr

Shivam Singh, Aditya Yadavalli, Catherine Arnett et al. · 0 citations
Aug 2026

Using accent variability to probe the performance of the Whisper automatic speech recognition system

Automatic speech recognition (ASR) systems often achieve high accuracy for native speech, yet remain less reliable for non-native (L2) accented speech. This gap raises a question about why ASR performs so well on L1 speech. When acoustic cues diverge from expectation, does ASR accommodate L2 speech acoustics, or does it rely on semantic predictability to infer likely words? To separate acoustic sensitivity from semantic inference, this study uses Whisper to probe how model size, semantic context, and talker-level variability influence transcription accuracy under accent-related variability. Five Whisper models (tiny, base, small, medium, and large-v3) were used to transcribe 200 read sentences, half high-predictability and half low-predictability, produced by 24 L1 Mandarin speakers of English and 24 L1 American-accented English speakers. We characterize each talker’s vocabulary knowledge, accent exposure, and perceived accentedness. Mixed-effects analyses of transcription accuracy will test three predictions. First, increasing model size will predict higher accuracy for both L1 and L2 speech, with an outsized effect for L2 speakers. Second, high-predictability sentence context will predict higher target-word accuracy, with a larger benefit for L2 speech. Third, continuous measures of lexical proficiency, prior accent exposure, and perceived accentedness will explain transcription accuracy beyond a categorical L1–L2 distinction.

Yuanrong Shen, Oishani Bandopadhyay, Sarah C. Creel · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.