Skip to content
Book Open access

The Alignment Gap: A Benchmark Demonstrating the Lack of Cross-Lingual Mapping in Dialect-Specialized Language Models - The Case of Ehugbo

Jul 2026 · Annual International ACM SIGIR Conference on Research and Development in Information Retrieval · pp. 3152-3158 · 0 citations · 26 references
Computer Science

TL;DR

This work introduces the first information retrieval benchmark resource for Ehugbo, a high-quality parallel multimodal corpus constructed from a high-quality parallel multimodal corpus that reveals the ''Alignment Gap'', where African-centric foundation models that excel at linguistic familiarity with Igbo achieve <5% retrieval accuracy while global models like LaBSE achieve 85% despite less exposure to the language family.

Abstract

Cross-lingual information retrieval (CLIR) for low-resource dialects remains underexplored, despite millions of speakers worldwide. This work addresses a critical gap by introducing the first information retrieval (IR) benchmark resource for Ehugbo (the Afikpo dialect of Igbo with \textasciitilde 150,000 speakers in Nigeria), constructed from a high-quality parallel multimodal corpus: 1 hour of transcribed Ehugbo Bible audio aligned with standard English translations. This parallel design enables rigorous evaluation of retrieval models across language pairs. Our benchmark reveals a surprising and counterintuitive finding: the ''Alignment Gap'', where African-centric foundation models (Serengeti, Afro-XLMR, AfriBERTA) that excel at linguistic familiarity with Igbo achieve <5% retrieval accuracy, while global models like LaBSE achieve 85% despite less exposure to the language family. Through diagnostic analysis (t-SNE visualizations, tokenization studies, error patterns), we show that regional models lack cross-lingual alignment bridges despite deep language understanding, while global models achieve language invariance through explicit translation supervision. This finding has immediate implications for the design of multilingual systems: pre-training diversity alone is insufficient for dialectal IR. We release our Ehugbo corpus, results, and evaluation splits on GitHub to enable future work on dialect-specific fine-tuning and alignment strategies for African languages.

Read PDF

Similar papers

Preprint Aug 2026

Interpretable Cross-Lingual Alignment in Small Language Models: Probing Cultural and Pragmatic Reasoning in Japanese-English Bilingual LLMs

J-PragEval-v0 is introduced, a minimal-pair benchmark isolating four such phenomena from surface fluency, and Pragmatic Representation Steering is specified, a parameter-free inference-time method that edits residual-stream activations along the class-mean-difference directions probing identifies.

F. Braun · 1 citation
#natural language process... Preprint Jul 2026

Viveka-Insight: a cross-lingual concept graph and citation-grounded retrieval resource over the complete works of Swami Vivekananda in English and Bengali

Classical philosophical corpora pose three compounding challenges for language resources: they exist in several languages without parallel alignment, their vocabulary is remote from that of contemporary readers, and generated text over culturally sensitive material must be verifiably grounded. We present Viveka-Insight...

Tamal Maharaj · 0 citations
Open access Aug 2026

Using Large Language Models in Formalizing Classical Linguistic Descriptions: A Case Study in Middle Indo-Aryan Sound Change

We present a systematic evaluation of Large Language Models (LLMs) in translating classical descriptions of phonological and morphophonological change from Old Indo-Aryan (Sanskrit) to Middle Indo-Aryan (MIA) into the standard notation of modern historical linguistics. Drawing on Vararuci’s Prākṛta Prakāśa (c. 4th ce...

V.S.D.S.Mahesh Akavarapu, Chinmay Dharurkar, Johannes Dellert et al. · 0 citations
Open access Aug 2026

Monolingual anchoring for low-resource cross-lingual semantic alignment: A case study on uyghur

Low-resource languages remain challenging for cross-lingual semantic alignment because of limited parallel corpora. In addition, conventional symmetric alignment may distort the semantic space of a high-resource language through noisy low-resource updates. To address this issue, we propose Monolingual Anchoring for Cro...

Ruohao Yan, Huaping Zhang, Yuwen Niu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.