Skip to content
Open access

Kannada Named Entity Recognition Using Deep Learning Techniques

2026 · International journal of research and scientific innovation · 0 citations

TL;DR

This paper proposes to use mBERT IndicBERT and XLM-RoBERTa language models pretrained on target and other related Indic language corpora and further fine-tune these models for the NER task.

Abstract

Named Entity Recognition (NER) is a natural language processing task concerned with identifying mentions of named entities and classifying them according to a predefined set of categories. Despite the success of NER in domains, where such data is abundant it remains a formidable challenge for low-resource languages such as Kannada. In this paper we discuss the possible ways to approach NER for the Kannada language. We explore various research directions including rule-based methods statistical machine learning neural networks and transformers based tagging methodologies. We highlight the various challenges in achieving NER for such a language and propose a transformer based contextual tagging framework for labelling sequences. We propose to use mBERT IndicBERT and XLM-RoBERTa language models pretrained on target and other related Indic language corpora and further fine-tune these models for the NER task. We discuss various aspects for experimentation including data collection labelling data preparation methods data-splits evaluation metrics comparison with other models hyper parameter tuning entity-wise analysis and error analysis.

Read PDF

Similar papers

Open access Aug 2026

NerAxom: a BIO-tagged NER dataset and hybrid neural-rule framework for Assamese

The NerAxom dataset is presented, a BIO-tagged NER dataset for Assamese comprising 4,173 sentences and 106,046 tokens annotated across seven entity categories, and a set of language-specific post-processing rules based on morphological suffixes and keyword cues are introduced.

Punam Sarmah, M. Lahkar, Shobhanjana Kalita et al. · 0 citations
Open access Jul 2026

PAP_NER: A large-scale vietnamese administrative named entity recognition corpus and hybrid deep learning architecture

PAP_NER is presented, the first large-scale, gold-standard Vietnamese administrative NER corpus comprising 162,801 sentences with 205,807 entity annotations across five entity types critical for e-Government workflows, with implications for low-resource language NLP research.

Dinh-Dien La, Tien-Bang Tran, Ngoc-Huy Du et al. · 0 citations
Open access Aug 2026

Ensemble-Based Approach for Amazigh POS Tagging: Leveraging Multiple Models for Enhanced Performance in Low-Resource Language Processing

The results show that hybrid ensemble methods can improve token-level accuracy in low-resource POS tagging, while also revealing a trade-off between frequent-tag accuracy and rare-tag robustness.

Abdelouahed Moussaoui, Nor-Eddine Azalmad, Said Bahassine et al. · 0 citations
Preprint Aug 2026

Unsupervised Multidomain Approaches to Named Entity Recognition with Small Datasets

This study applies an unsupervised pre-training approach to precondition and identify entities without annotated datasets, then applies transfer learning models to different simulated limited datasets for a named entity recognition task.

Israel Fianyi, James Montgomery, Soon-Jeong Yeom · 0 citations
#artificial intelligence Preprint Sep 2026

NE-R1: Enhancing Named Entity Recognition Model via Reinforcement Learning

Named Entity Recognition (NER) has achieved substantial progress since the advent of large language models (LLMs). Nevertheless, the recognition of long-tail and domain-specific entities remains challenging due to the deficiency in parametric knowledge. Retrieval-augmented generation (RAG) offers a promising remedy by...

Mei-Xuan Chen, He-Han Li, Rui-Zhi Zhao et al. · 0 citations
Open access 2026

Leveraging Task-Adaptive Continual Pre-Training to Enhance the Classification Ability of Language Models

This study presents a systematic empirical investigation of task-adaptive continual pre-training (TAPT), introduced by Gururangan et al., for Turkish language understanding, with a particular focus on the effect of the masked-language-modeling rate.

Murat Aydoğan, Savas Yildirim, Tuǧba Dalyan · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.