Aug 2026· WORLD JOURNAL OF INNOVATION AND MODERN TECHNOLOGY· 0 citations
TL;DR
A transferable curation framework for endangered languages, the first sizable Kalabari-English parallel corpus, and baseline experiments that reveal both the promise and the hallucination pitfalls of training on highly constrained, domain-specific data are contributed.
Abstract
The development of Natural Language Processing (NLP) tools for endangered and low
resource languages is fundamentally hindered by the scarcity of high-quality parallel data.
Data for languages like Kalabari is not only scarce but often noisy, inconsistently digitized,
and orthographically unstandardized. While prior work has leveraged religious texts for
corpus creation, the specific challenges of extracting and normalizing morphologically rich
languages with complex diacritics remain underexplored. This paper addresses this gap by
introducing a reproducible, modular curation methodology tailored for such languages. We
document a six-step pipeline that transforms raw digital texts—sourced from the Kalabari
Bible (FiaFia Biabulu) and instructional literature (Kalabari Lingua)—into a clean, verse
aligned, 10,222-pair parallel corpus. We demonstrate that enforcing Normalization Form C
(NFC) is critical for preserving sub-dot diacritics, and we validate the corpus by training a
baseline Transformer NMT system. Our contributions are threefold: (1) a transferable curation
framework for endangered languages, (2) the first sizable Kalabari-English parallel corpus,
and (3) baseline experiments that reveal both the promise and the hallucination pitfalls of
training on highly constrained, domain-specific data.
The results demonstrate a reproducible, CPU-centric pipeline, proving that the lack of specialized GPU infrastructure is not an insurmountable obstacle for digital language preservation and baseline NMT development.
O. T. Olise· International Journal of Com...· 0 citations
A corpus containing both manually curated and synthetically generated sentences for low-resource Indian languages, such as Maithili is contributed and it is demonstrated that, even with a smaller corpus size, high-quality, task-specific data significantly enhance translation accuracy for low-resource Indian languages,...
Kamanksha Prasad Dubey, C. Maurya, Kumar Padmanabh· International Conference on...· 0 citations
A systematic review of stemming techniques across Afro-Asiatic, Indo-Aryan, Turkic, and Uralic language families, with explicit acknowledgment of coverage limitation, points out the strengths and weaknesses of current approaches and provides insights into potential areas for future investigation and developments in lan...
Sanjiban Sekhar Roy, Mohd Anas, Saravanakumar Kandasamy· Frontiers in Artificial Inte...· 0 citations
Dasinya, the first openly released Badini Kurdish text corpus, together with the Badini Dialect Processing Toolkit (BDPT) — the first NLP preprocessing pipeline built specifically for this dialect are presented.
V. A. Saeed, Karwan Jacksi· Data in Brief· 0 citations