Skip to content
Open access

NerAxom: a BIO-tagged NER dataset and hybrid neural-rule framework for Assamese

Aug 2026 · Engineering Research Express · Vol 8 · 0 citations · 49 references
Physics

TL;DR

The NerAxom dataset is presented, a BIO-tagged NER dataset for Assamese comprising 4,173 sentences and 106,046 tokens annotated across seven entity categories, and a set of language-specific post-processing rules based on morphological suffixes and keyword cues are introduced.

Abstract

Named entity recognition (NER) in low-resource, morphologically rich languages such as Assamese (ISO 639-3: asm) remains a significant challenge due to the scarcity of annotated corpora and the limited applicability of models designed for resource-rich languages. Existing Assamese NER resources suffer from critical limitations: WikiAnn provides broad language coverage but insufficient data volume for neural model training; AsNER, while a gold-standard corpus, supports only five entity categories and lacks a formal tagging scheme, restricting its utility for downstream tasks such as relation extraction and information retrieval. Furthermore, prior Assamese NER systems have relied predominantly on traditional tagging approaches and classical machine learning methods, with limited exploration of modern pre-trained language models and linguistically motivated post-processing strategies. To address these gaps, we present NerAxom, a BIO-tagged NER dataset for Assamese comprising 4,173 sentences and 106,046 tokens annotated across seven entity categories: Person (PER), Location (LOC), Organization (ORG), Date (DATE), Work_of_Art (WOA), Occupation (OCC), and Number (NUM). The dataset was independently annotated by two trained native speakers, achieving an inter-annotator agreement of κ=0.82 (Cohen’s Kappa), with disagreements resolved through expert linguist adjudication. We evaluate NerAxom using two modeling paradigms: (i) a BiLSTM–CRF model with an attention mechanism, tested with FastText, BERT, and MuRIL embeddings; and (ii) direct fine-tuning of the MuRIL transformer as a token classifier. Among embedding-based models, MuRIL yields the highest F1-score of 68%, outperforming FastText (62%) and BERT (64%). Fine-tuning MuRIL directly as a token classifier achieves an F1-score of 70%, establishing a competitive transformer baseline. To address entity misclassifications arising from Assamese morphological complexity, we further introduce a set of language-specific post-processing rules based on morphological suffixes and keyword cues. The hybrid system combining MuRIL embeddings in the BiLSTM–CRF+Attention architecture with these linguistic rules achieves an F1-score of 71% and an overall accuracy of 82% on Assamese Wikipedia biographical text, competitive with the fine-tuned MuRIL transformer (70% F1) and demonstrating that linguistically informed post-processing provides complementary gains over embedding-based neural baselines. The rule component yields the largest category-wise gains for LOC (+20 F1), WOA (+11 F1), and ORG (+9 F1). The NerAxom dataset is publicly available to support further NER research in Assamese and related low-resource Indic languages.

Read PDF

Similar papers

Open access 2026

Kannada Named Entity Recognition Using Deep Learning Techniques

This paper proposes to use mBERT IndicBERT and XLM-RoBERTa language models pretrained on target and other related Indic language corpora and further fine-tune these models for the NER task.

P. M, H. G, S. N · 0 citations
Open access 2026

Legal NER: Evaluating the Impact of LLM-Generated Annotations on NER Performance for Administrative Decisions

This study investigates the use of LLM-generated annotations to expand the training set for supervised NER models applied to sentences from Dutch administrative decisions as a low-resource domain and language and indicates that LLMs can accurately generate annotations for legal entities that are explicitly defined in l...

H. Nan, Samaneh Khoshrou, Johan Wolswinkel · 0 citations
Open access Aug 2026

Ensemble-Based Approach for Amazigh POS Tagging: Leveraging Multiple Models for Enhanced Performance in Low-Resource Language Processing

The results show that hybrid ensemble methods can improve token-level accuracy in low-resource POS tagging, while also revealing a trade-off between frequent-tag accuracy and rare-tag robustness.

Abdelouahed Moussaoui, Nor-Eddine Azalmad, Said Bahassine et al. · 0 citations
Open access Jul 2026

PAP_NER: A large-scale vietnamese administrative named entity recognition corpus and hybrid deep learning architecture

PAP_NER is presented, the first large-scale, gold-standard Vietnamese administrative NER corpus comprising 162,801 sentences with 205,807 entity annotations across five entity types critical for e-Government workflows, with implications for low-resource language NLP research.

Dinh-Dien La, Tien-Bang Tran, Ngoc-Huy Du et al. · 0 citations
Open access Aug 2026

Optimizing sample selection for large language model-based entity matching using AssistEM

AssistEM, a framework for efficient LLM adaptation to EM via principled data selection, demonstrates that selective fine-tuning not only accelerates adaptation but also improves training efficiency (requiring fewer GPU hours), enabling open-source LLMs to rival–and in some cases outperform–closed-source models.

John Bosco Mugeni, Steven J. Lynden, Toshiyuki Amagasa et al. · 1 citation
Open access Aug 2026

Biomedical Text Mining and Information Extraction Using Prompt-Enhanced and LoRA-Adapted Large Language Models

Biomedical named entity recognition (NER) and relation extraction (RE) remain challenging because biomedical texts contain ambiguous abbreviations, complex entity boundaries, domain-specific terminology, and implicit relations. This study proposes a prompt-enhanced and QLoRA-adapted large language model framework for b...

Feng Yan, De-Quan Zheng, Feng Yu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.