Jul 2026· Annual International ACM SIGIR Conference on Research and Development in Information Retrieval· pp. 3152-3158· 0 citations· 26 references
Computer Science
TL;DR
This work introduces the first information retrieval benchmark resource for Ehugbo, a high-quality parallel multimodal corpus constructed from a high-quality parallel multimodal corpus that reveals the ''Alignment Gap'', where African-centric foundation models that excel at linguistic familiarity with Igbo achieve <5% retrieval accuracy while global models like LaBSE achieve 85% despite less exposure to the language family.
Abstract
Cross-lingual information retrieval (CLIR) for low-resource dialects remains underexplored, despite millions of speakers worldwide. This work addresses a critical gap by introducing the first information retrieval (IR) benchmark resource for Ehugbo (the Afikpo dialect of Igbo with \textasciitilde 150,000 speakers in Nigeria), constructed from a high-quality parallel multimodal corpus: 1 hour of transcribed Ehugbo Bible audio aligned with standard English translations. This parallel design enables rigorous evaluation of retrieval models across language pairs. Our benchmark reveals a surprising and counterintuitive finding: the ''Alignment Gap'', where African-centric foundation models (Serengeti, Afro-XLMR, AfriBERTA) that excel at linguistic familiarity with Igbo achieve <5% retrieval accuracy, while global models like LaBSE achieve 85% despite less exposure to the language family. Through diagnostic analysis (t-SNE visualizations, tokenization studies, error patterns), we show that regional models lack cross-lingual alignment bridges despite deep language understanding, while global models achieve language invariance through explicit translation supervision. This finding has immediate implications for the design of multilingual systems: pre-training diversity alone is insufficient for dialectal IR. We release our Ehugbo corpus, results, and evaluation splits on GitHub to enable future work on dialect-specific fine-tuning and alignment strategies for African languages.
Within low resource domains, results identify the model's parsing of information and subsequent reasoning as the source of reasoning failure, rather than corpus contents, rather than corpus contents.
J-PragEval-v0 is introduced, a minimal-pair benchmark isolating four such phenomena from surface fluency, and Pragmatic Representation Steering is specified, a parameter-free inference-time method that edits residual-stream activations along the class-mean-difference directions probing identifies.
The Cross-Lingual Comprehension Gap (CLCG) is defined as the reduction in response quality when the same content and question are presented in a target language rather than in English.
Classical philosophical corpora pose three compounding challenges for language resources: they exist in several languages without parallel alignment, their vocabulary is remote from that of contemporary readers, and generated text over culturally sensitive material must be verifiably grounded. We present Viveka-Insight...
We present a systematic evaluation of Large Language Models (LLMs) in translating classical descriptions of phonological and morphophonological change from Old Indo-Aryan (Sanskrit) to Middle Indo-Aryan (MIA) into the standard notation of modern historical linguistics. Drawing on Vararuci’s Prākṛta Prakāśa (c. 4th ce...
V.S.D.S.Mahesh Akavarapu, Chinmay Dharurkar, Johannes Dellert et al.· Computational Linguistics· 0 citations
Low-resource languages remain challenging for cross-lingual semantic alignment because of limited parallel corpora. In addition, conventional symmetric alignment may distort the semantic space of a high-resource language through noisy low-resource updates. To address this issue, we propose Monolingual Anchoring for Cro...
Ruohao Yan, Huaping Zhang, Yuwen Niu et al.· Journal of King Saud Univers...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.