Skip to content
Preprint

Bulbul: A Dataset for Dialectal Arabic Speech Recognition

Aug 2026 · 0 citations · 34 references
Computer Science

TL;DR

BUL, a multi-dialect Arabic ASR dataset collected from 275 speakers in 11 Arab countries, includes structured dialect and sub-dialect coverage, as well as recordings of classical Arabic and modern standard Arabic spoken by participants in their native dialectal accents to support accent-aware modeling.

Abstract

Arabic automatic speech recognition (ASR) faces unique challenges due to diglossia, extensive regional dialect variation, and limited speech resources. Existing speech datasets often focus on single dialects or large-scale broadcast/web data, leading to trade-offs between linguistic diversity and annotation quality. We present BULBUL, a multi-dialect Arabic ASR dataset collected from 275 speakers in 11 Arab countries. BULBUL includes structured dialect and sub-dialect coverage, as well as recordings of classical Arabic and modern standard Arabic spoken by participants in their native dialectal accents to support accent-aware modeling. The quality of the recordings was ensured through a two-level human verification process. We further benchmark a range of recent ASR systems, establishing strong baselines for modern dialectal and accented Arabic ASR.

View source

Similar papers

Review Aug 2026

Romanized Arabic Across Dialects: Views, Usage Patterns, and Linguistic Variation

Arabizi refers to Arabic written in Latin script. Although previous studies have shown that the prevalence and usage of Arabizi vary by factors such as region and age group, most NLP research on Arabic texts treats it as a temporary phenomenon resulting from limited technological support for the Arabic script. In this work, we engage with Arabic speakers to collect insights on their perceptions and usage of Arabizi. We further examine writing norms among speakers of different dialects, focusing on Algerian, Egyptian, Lebanese, Moroccan, and Tunisian Arabic. To this end, we release two resources. First, a character-level alignment of Arabic words to study inter- and intra-dialectal variation across these five dialects, based on words transliterated by survey participants, finding systematic intra-dialectal regularity and inter-dialectal variation. Second, to study Arabic speakers'ability to identify this stylistic variation at the sentence-level, we build a manually curated parallel corpus of sentences written in Arabic script alongside multiple Arabizi transliterations, collected from speakers of the same five dialects. Our study presents the largest human-centered, cross-dialectal study of Arabizi's perceptions and practices to date.

Amr Keleg, Ahmed Amine Ben Abdallah, Taha Yassine et al. · 0 citations

Aladdin­FTI @ AMIYA Three Wishes for Arabic NLP: Fidelity, Diglossia, and Multidialectal Generation

Aladdin­FTI, the proposed system is designed to both generate and translate dialectal Arabic (DA), and supports text generation in Moroccan, Egyptian, Palestinian, Syrian, and Saudi dialects, as well as bidirectional translation between these dialects, Modern Standard Arabic, and English.

Jonathan Mutal, Perla Al Almaoui, Simon Hengchen et al. · 0 citations
Open access 2026

Exploring Arabic Hate Speech Detection: Dialect Variations and Multimodal Approaches

With the growing use of social networks and online platforms, hate speech and racist content detection is becoming increasingly challenging, especially in the Arab world, which is characterized by a significant diversity of complex dialects. This article highlights the importance of implementing detection systems capable of identifying hate speech, based on a critical analysis of scientific literature published between 2020 and 2025. This paper highlights the diversity of Arabic dialects and their linguistic complexity, which make them difficult to analyze using a single model. As a result, the utility of using multimodal approaches for detecting hate speech is emphasized. The article presents this approach and evaluates the effectiveness of several artificial intelligence-based methods, implementing methodological innovations and identifying gaps in current research, as well as addressing key challenges such as the limited consideration of Maghrebi dialects, the lack of datasets for these dialects, especially Moroccan Darija.

M. Hassane, M. Azrour, A. Hessane et al. · 0 citations
Open access Jul 2026

A Nepali-Accented English Evaluation Dataset for Automatic Speech Recognition

A Nepali-accented English evaluation dataset designed to support robust ASR benchmarking under accent mismatch is presented, positioning the corpus as a practical held-out resource for evaluating accent robustness and out-of-distribution generalization on Nepali-accented English.

Santosh Dahal, K. Dahal · 0 citations
#artificial intelligence Preprint Sep 2026

EDRAC: Benchmarking Arabic Dialect Reading Comprehension

This work introduces EDRAC, the first large-scale benchmark for dialectal Arabic machine reading comprehension (MRC) and generative QA, covering five major dialects: Egyptian, Moroccan, Emirati, Syrian, and Saudi Arabic, and benchmarks Arabic-centric and multilingual LLMs on EDRAC using lexical and semantic metrics.

Noor Abo Mokh, K. Chirkunov, Teresa Lynn et al. · 0 citations
Jul 2026

Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages

In this work, we introduce Indic DiarBench, a speaker diarization and ASR benchmark dataset spanning all 22 scheduled languages of India. This corpus comprises approximately 108 hours of natural multi-speaker audio from near-field meetings, far-field recordings, and in-the-wild audios. All annotations are human-corrected with time-aligned speaker attributed transcriptions. The dataset captures conversational nuance prevalent in Indian speech, such as English code-mixing, dialectal variation, and frequent speaker overlap. To establish a baseline for joint ASR and diarization capabilities we evaluate leading systems including commercial speech APIs and multimodal large language models. Indic DiarBench is released as an open-access resource to advance inclusive, multilingual speech technology research for Indian languages.

Deovrat Mehendale, Aditya Mehndiratta, Dhruv Rathi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.