Skip to content
Preprint

Unsupervised Multidomain Approaches to Named Entity Recognition with Small Datasets

Aug 2026 · 0 citations · 33 references
Computer Science

TL;DR

This study applies an unsupervised pre-training approach to precondition and identify entities without annotated datasets, then applies transfer learning models to different simulated limited datasets for a named entity recognition task.

Abstract

This paper explores the challenges and the methodologies associated with learning quality representations in scenarios with unlabelled small or limited datasets for downstream information extraction task (Multidomain Named Entity Recognition (NER). The study adopts a Transfer Learning on small datasets. Traditional NER systems often rely on large, labelled data, which is impractical for many domains. This study, therefore, applies an unsupervised pre-training approach to precondition and identify entities without annotated datasets, then applies transfer learning models to different simulated limited datasets for a named entity recognition task. Entity Recognition (NER) is essential in natural language processing (NLP), it identifies and classifies related entities within the text. This study addresses the complexities of domain variability, data sparsity, and overfitting and investigates innovative approaches such as data augmentation, few-shot learning, and domain adversarial training. Integrating these techniques promises to enhance the performance and generalizability of NER systems across diverse and resource-constrained domains, paving the way for more efficient and adaptable NLP applications.

View source

Similar papers

Open access 2026

Kannada Named Entity Recognition Using Deep Learning Techniques

This paper proposes to use mBERT IndicBERT and XLM-RoBERTa language models pretrained on target and other related Indic language corpora and further fine-tune these models for the NER task.

P. M, H. G, S. N · 0 citations
Open access Aug 2026

NerAxom: a BIO-tagged NER dataset and hybrid neural-rule framework for Assamese

The NerAxom dataset is presented, a BIO-tagged NER dataset for Assamese comprising 4,173 sentences and 106,046 tokens annotated across seven entity categories, and a set of language-specific post-processing rules based on morphological suffixes and keyword cues are introduced.

Punam Sarmah, M. Lahkar, Shobhanjana Kalita et al. · 0 citations
Open access Sep 2026

A Lightweight Framework for Medical Named Entity Recognition and Relation Extraction.

INTRODUCTION Clinical text in electronic health records contains valuable information for medical research. Named entity recognition (NER) and relation extraction (RE) are key tasks for structuring this information, but existing deep learning approaches are often computationally expensive. METHODS We develop a lightw...

Sihan Wu, Susanna Gordon, M. Boeker et al. · 0 citations

Addressing Span Imbalance and Semantic Complexity in Nested Medical Named Entity Recognition

The proposed AGPNer integrates a heterogeneous dependency fusion encoder, which reconstructs masked entities to enhance token representations and fine-tunes a hybrid dependency modeling block to learn domain-specific patterns in medical texts; an imbalance-adaptive span decoder, which decouples entity and non-entity sp...

Yuling Li, Yang-Juan Hu, Yi-Ming Bao et al. · 0 citations
#small language model Open access Oct 2026

Named Entity Recognition Across Datasets and Domains: Resources and Models for Galician

Automatic named entity recognition (NER) is essential for many natural language processing applications, particularly in low-resource languages and those with corpora restricted to a single domain, where the lack of diverse data may hinder cross-domain generalisation. In this context, we attempt to contribute to the...

A. Sarymsakova, Ettore Mariotti, Helena Pérez Puente et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.