Skip to content
Open access

A Nepali-Accented English Evaluation Dataset for Automatic Speech Recognition

Jul 2026 · Everest Advances in Science and Technology · 0 citations · 7 references

TL;DR

A Nepali-accented English evaluation dataset designed to support robust ASR benchmarking under accent mismatch is presented, positioning the corpus as a practical held-out resource for evaluating accent robustness and out-of-distribution generalization on Nepali-accented English.

Abstract

Automatic speech recognition (ASR) systems perform strongly on native-English benchmarks, yet their accuracy degrades sharply when the input speech comes from under represented non-native accents. Nepali-accented English is particularly under-served: existing resources either focus on native Nepali speech, cover broader multi-accent settings without dedicated Nepali evaluation, or provide only limited Nepali-accent coverage. This paper presents a Nepali-accented English evaluation dataset designed to support robust ASR benchmarking under accent mismatch. The corpus was collected through a web-based platform that did not collect directly identifying metadata and contains recordings from 57 speakers. Each session follows a fixed 22-prompt protocol consisting of 11 phonetic prompts, 10 domain prompts, and 1 spontaneous prompt, providing complementary coverage of pronunciation, topical vocabulary, and natural speaking style. In addition to transcribed speech, the dataset includes participant metadata for coarse exploratory subgroup analysis and speaker-level manual recording-quality labels. Manual quality assessment shows that 50.9% of sessions are clean and 42.1% contain only mild noise. As a descriptive reference, open-source ASR baselines are substantially worse on this corpus than the corresponding LibriSpeech test-clean values reported in official model cards, reaching 38.15–55.00% WER on the collected set versus reported 2–4% WER on LibriSpeech test-clean. These baseline results position the corpus as a practical held-out resource for evaluating accent robustness and out-of-distribution generalization on Nepali-accented English.

Read PDF

Similar papers

Preprint Sep 2026

Phoneme-Aware Pronunciation Representations for L2-English L1-Background Accent Identification

We study speaker-disjoint accent identification for L2 English, where the goal is to predict a speaker's first-language (L1) background from English pronunciation. Most existing systems classify accents using a single utterance-level representation, but such global representations can obscure pronunciation cues that depend on specific English phonemes. We propose a transcript-assisted model that makes phoneme information explicit during accent identification. Instead of representing an utterance only as a global speech embedding, we represent it as a sequence of pronunciation units, each combining acoustic evidence from a spoken segment with the aligned English phoneme for that segment. A frozen speech encoder provides the acoustic features, while the transcript is used only to obtain phoneme-level forced alignments. No word-level or sentence-level text representation is passed to the accent classifier. Under a four-fold speaker-disjoint protocol on L2-ARCTIC, our model achieves 81.41% accuracy and 81.21% macro-F1, the highest mean performance among the evaluated systems. Diagnostic ablations support the importance of phoneme-aligned token construction, while a Whisper-based ablation shows an additional gain from phoneme information.

Yang-Yang Qu, Massimiliano Todisco, Nicholas W. D. Evans · 0 citations
#small language model Preprint Aug 2026

Fine-Tuning Whisper for Automatic Speech Recognition in Baniwa: A Preliminary Study

A preliminary study on the adaptation of Whisper for Automatic Speech Recognition in Baniwa, an indigenous Arawakan language spoken in Brazil, Colombia, and Venezuela, demonstrating that multilingual foundation models can be successfully adapted to extremely low-resource indigenous languages.

Leonardo Duart, T. Fonseca, T. Chacon · 0 citations
Preprint Aug 2026

SraVaani 1.0: Scaling Inclusive Speech Recognition for Indic Languages

SraVaani-1.0 achieves the lowest word error rate (WER) on a large number of language-dataset pairs while remaining competitive with the best-performing systems on high resource while being assessed exclusively on the VAANI benchmark.

Sujith Pulikodan, A. Basu, J. Pavankumar et al. · 1 citation · ⚡1

Neyshekar: An Open Persian Read-Speech Corpus for Automatic Speech Recognition

Neyshekar is presented as an open Persian read-speech corpus designed for coverage of both formal and informal language, named entities, and longer utterances. In version 6, 62,279 validated recordings totalling 99.02 hours are provided from 190 contributors, with 34,541 distinct recorded prompts. The prompt pool was assembled from human-written material, contextualised homographs, and reviewed language-model-generated text. Text entries were normalised with the shekar library, which supports both formal and informal Persian, and every submitted recording was reviewed against a common validation rubric. About 24% of released clips are classified as informal by an automatic classifier; these register labels are not human-validated. Item-level rater labels are provided for reproducible agreement estimation, opaque per-clip contributor identifiers make the speaker-disjoint partitioning auditable and support contributor-clustered uncertainty estimates, and a text-disjoint test subset is included for evaluation beyond previously seen prompts. Per-contributor recording load and reference-free signal quality are characterised for every released clip. Corpus characteristics are compared with Persian Common Voice under shared processing. Utility is assessed through two ASR architectures, three optimisation seeds, WER and CER, and independent evaluation on the public PSRB sample. Against duration-matched Common Voice training at approximately 32 hours, in-domain WER is reduced by 9.5 points for Whisper and 11.6 points for XLS-R, and by approximately eight points for both architectures on the independent PSRB sample. Transfer and mixture benefits are not consistently observed across architectures and training budgets. The corpus is released under CC0; code and data are made available through the project repository at https://github.com/amirivojdan/neyshekar.

Ahmad Amirivojdan, Farzad Nadiri, Abolfazl Alizadeh et al. · 0 citations
Aug 2026

Using accent variability to probe the performance of the Whisper automatic speech recognition system

Automatic speech recognition (ASR) systems often achieve high accuracy for native speech, yet remain less reliable for non-native (L2) accented speech. This gap raises a question about why ASR performs so well on L1 speech. When acoustic cues diverge from expectation, does ASR accommodate L2 speech acoustics, or does it rely on semantic predictability to infer likely words? To separate acoustic sensitivity from semantic inference, this study uses Whisper to probe how model size, semantic context, and talker-level variability influence transcription accuracy under accent-related variability. Five Whisper models (tiny, base, small, medium, and large-v3) were used to transcribe 200 read sentences, half high-predictability and half low-predictability, produced by 24 L1 Mandarin speakers of English and 24 L1 American-accented English speakers. We characterize each talker’s vocabulary knowledge, accent exposure, and perceived accentedness. Mixed-effects analyses of transcription accuracy will test three predictions. First, increasing model size will predict higher accuracy for both L1 and L2 speech, with an outsized effect for L2 speakers. Second, high-predictability sentence context will predict higher target-word accuracy, with a larger benefit for L2 speech. Third, continuous measures of lexical proficiency, prior accent exposure, and perceived accentedness will explain transcription accuracy beyond a categorical L1–L2 distinction.

Yuanrong Shen, Oishani Bandopadhyay, Sarah C. Creel · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.