Skip to content
Open access

Dasinya: A dataset and preprocessing pipeline for the Badini Kurdish dialect

Aug 2026 · Data in Brief · Vol 68, pp. 113129 · 0 citations · 19 references
Medicine

TL;DR

Dasinya, the first openly released Badini Kurdish text corpus, together with the Badini Dialect Processing Toolkit (BDPT) — the first NLP preprocessing pipeline built specifically for this dialect are presented.

Abstract

Despite notable advances in NLP for low-resource languages, dialects written in non-Latin scripts — especially those without standardised digital resources — continue to receive insufficient attention. The Badini Kurdish dialect, written in Perso-Arabic script and spoken across the Kurdistan Region of Iraq, lacks any dedicated NLP preprocessing pipeline in the literature. This paper presents Dasinya, the first openly released Badini Kurdish text corpus, together with the Badini Dialect Processing Toolkit (BDPT) — the first NLP preprocessing pipeline built specifically for this dialect. Dasinya contains 107 files, 87,545 sentences, and roughly 1.28 million words, spanning two source types and nine sub-genres: six book sub-genres (literary, historical, cultural, scientific, psychological, and educational) and three non-book sub-genres (news, academic/linguistic, and broadcast). These files were drawn from two repositories: the Badirxanian Public Library in Duhok (98 books) and the Semantic Web Lab at the University of Zakho (4 books and 5 non-book files). BDPT is structured as a six-stage pipeline that handles validation, Unicode normalisation — enforcing Badini-specific character mappings and preserving the Zero-Width Non-Joiner as a morpheme boundary marker — structural cleaning, deep cleaning, text preprocessing, and sentence segmentation. Orthographic problems particular to Badini Kurdish are addressed at each stage; none of these are handled by existing Arabic, Persian, or Sorani tools. At Stage 6, sentence segmentation reaches a specialist-validated Acc% of 91.4% across all 107 files, measured under a two-metric framework (OK% and Acc%) with genre-sensitive thresholds that reflect genuine linguistic differences between books, news, broadcast, and academic texts. Both the Dasinya corpus and the BDPT pipeline are openly available at https://doi.org/10.5281/zenodo.20729060 and are intended to underpin future Badini Kurdish NLP work, including tokenisation, part-of-speech tagging, stemming, and stopword removal.

Read PDF

Similar papers

#natural language process... Preprint Sep 2026

5-Dialects-BN: Unmasking the Impact of Transliteration on Bangla Dialectal LLMs

Large Language Models (LLMs) have achieved remarkable progress across natural language processing (NLP) tasks, yet their capabilities degrade sharply for low-resource languages and dialectally diverse settings. Bangla, the world's sixth most spoken language, exemplifies this gap: existing resources overwhelmingly targe...

Md Mahir Jawad, Galib Mahmud Jim, Rafid Ahmed et al. · 0 citations
Open access Aug 2026

Dataset Curation for Kalabari NMT System

A transferable curation framework for endangered languages, the first sizable Kalabari-English parallel corpus, and baseline experiments that reveal both the promise and the hallucination pitfalls of training on highly constrained, domain-specific data are contributed.

O. T. Olise · 0 citations
#natural language process... Preprint Oct 2026

BanglaDial-Abuse: A Corpus-Grounded Dataset for Regional Dialect Identification in Abusive Bangla Text

Regional linguistic variation remains an important challenge for Bangla natural language processing, particularly in informal and non-standard text. This paper introduces BanglaDial-Abuse, a balanced Bengali-script dataset developed for regional dialect identification in abusive and hostile Bangla text. The dataset con...

Hasin Almas Sifat · 0 citations
#machine learning Preprint Sep 2026

TutlAit v1: a crowdsourced Moroccan Tamazight speech dataset with Arabic transcriptions and regional accent labels

Tamazight (Amazigh) is, together with Arabic, one of the two official languages of Morocco, yet it remains severely under-resourced for speech technology: pub licly available labelled audio is scarce, generally lacks information on the regional variety spoken, and is often of uneven transcription quality. This article...

M. Chadi, Ezzahra Ait El Arbi, Ismail Khayoub et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.