Skip to content

DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation

Jul 2026 · arXiv.org · Vol abs/2607.07669 · 1 citation · ⚡ 1 influential · 45 references
Computer Science

TL;DR

DiaLLM is introduced, which continually pretrains three open-weight language model families on the International Corpus of English and applies implicit and explicit post-training paradigms, each combined with three model alignment strategies, giving the first controlled comparison of these components across Australian, Indian, and Northern British English.

Abstract

Large language models increasingly understand dialectal English, yet still produce only standard, US-leaning English, leaving dialectal generation, the harder half of the problem, largely unaddressed. We introduce DiaLLM, which continually pretrains three open-weight language model families on the International Corpus of English and applies implicit and explicit post-training paradigms, each combined with three model alignment strategies, giving the first controlled comparison of these components across Australian, Indian, and Northern British English. Our results reveal a robustness-generation gap: benchmarks are shaped by continual pretraining and SFT, while alignment visibly reshapes generation in ways benchmarks do not capture. Explicit variety-targeted adaptation produces output reliably recognised as dialectal and judged more dialectal than broad alignment, yet where human judgement was directly assessed, the method that most aggressively optimises the dialectal reward is not the one judged most dialectal. Independent linguistic analysis corroborates this reward-quality gap, most clearly on two of the three families. No single alignment method dominates, and closing the gap will require richer reward designs and continued investment in dialectal resources. We release all code, checkpoints, and preference datasets.

View source

Similar papers

Preprint Aug 2026

How Robust Are LLMs to Vietnamese Dialects?

These findings show that strong performance on Standard Vietnamese does not guarantee reliable behavior under meaning-preserving regional variation, and the first systematic evaluation of LLM robustness to Vietnamese dialect variation across multiple tasks is presented.

Minh Tran, C. Trinh, T. Lê et al. · 0 citations
Preprint Aug 2026

The Dialect Tax: Dialectal Biases Persist throughout the Language Modeling Pipeline

Systematic dialectal performance gaps in language models (LMs) are well documented, but the source of these disparities within the modern language modeling pipeline remains unclear. Our study traces this"dialect tax"across the natural language processing pipeline. Using parallel English dialect corpora that hold meanin...

Elle Michelle Yang, Mark Chen, Jerry Tworek et al. · 0 citations
Preprint Aug 2026

Interpretable Cross-Lingual Alignment in Small Language Models: Probing Cultural and Pragmatic Reasoning in Japanese-English Bilingual LLMs

J-PragEval-v0 is introduced, a minimal-pair benchmark isolating four such phenomena from surface fluency, and Pragmatic Representation Steering is specified, a parameter-free inference-time method that edits residual-stream activations along the class-mean-difference directions probing identifies.

F. Braun · 1 citation
#natural language process... Preprint Sep 2026

5-Dialects-BN: Unmasking the Impact of Transliteration on Bangla Dialectal LLMs

Large Language Models (LLMs) have achieved remarkable progress across natural language processing (NLP) tasks, yet their capabilities degrade sharply for low-resource languages and dialectally diverse settings. Bangla, the world's sixth most spoken language, exemplifies this gap: existing resources overwhelmingly targe...

Md Mahir Jawad, Galib Mahmud Jim, Rafid Ahmed et al. · 0 citations

Cultural Benchmarking of LLMs in MSA and Arabic Dialectal Dialogue

ArabCulture-Dialogue is introduced, a culturally grounded conversational dataset covering 13 Arabic-speaking countries, in both MSA and each country’s respective dialect, spanning 12 daily-life topics and 54 fine-grained subtopics, to address the performance gap between MSA and Arabic dialects.

M. Dehan, Aldry Kautsar, SaeedAlmheiri et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.