Skip to content
Preprint

AfriSwitch: A Benchmark for In-the-Wild African Code-Switched Speech Recognition

Aug 2026 · 0 citations · 40 references
Computer Science

TL;DR

AfriSwitch is presented, a 61.36-hour human-transcribed benchmark of in-the-wild code-switched speech spanning 16 African languages and language varieties, released with switch-level English span tags, perutterance Code-Mixing Index (CMI), and switch-point counts.

Abstract

Code-switching is pervasive in bilingual African conversation, yet most ASR systems assume monolingual input and are evaluated on curated monolingual benchmarks. We present AfriSwitch, a 61.36-hour human-transcribed benchmark of in-the-wild code-switched speech spanning 16 African languages and language varieties, released with switch-level English span tags, perutterance Code-Mixing Index (CMI), and switch-point counts. Corpus statistics show that mixing behaviour varies widely across African languages along two largely independent axes: how often speakers alternate, and how balanced the mixture is. No single scalar captures how code-switched a language is. Benchmarking five open and commercial multilingual ASR systems zero-shot yields word error rates far above published monolingual figures for the same languages, with the best system averaging 35.93% WER and no system falling below 24% on any language. Africa-targeted training, not model scale or nominal language coverage, best predicts performance.

View source

Similar papers

#natural language process... Preprint Sep 2026

Benchmarking Automatic Speech Recognition Tools for Iberian Languages

Comprehensive evaluations of automatic speech recognition (ASR) for Iberian languages remain limited, and low-resource languages, biases, and efficiency trade-offs are underexplored. We benchmark eleven systems, ten open-weight models and one commercial API, across five Iberian languages (Basque, Catalan, Galician, Por...

Fernando López, Pablo Gómez, David Solans et al. · 0 citations
#natural language process... Preprint Sep 2026

VietPrism: A large-scale Vietnamese speech and deepfake corpus with diverse dialects and code-switching

Vietnamese speech research is constrained by resources that isolate automatic speech recognition from speaker, dialect, code-switching, and deepfake analysis. We introduce VietPrism, an open, multi-domain corpus that brings these dimensions together at scale: 993.4 hours and 403,941 bona fide utterances from 1,262 veri...

M. Hoàng, Thai Le · 0 citations
#natural language process... Preprint Sep 2026

IndicFDB: Benchmarking Full-Duplex Voice Agents across Indian Languages

This work addresses three challenges: finding conversational events in multilingual speech, evaluating their timing without reliable word-level alignment, and judging responses across languages with IndicFDB, which extends it to ten languages spoken in India with 12,350 samples.

Rajarshi Roy, Shobhit Banga, Jonathan Raiman et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.