This study applies an unsupervised pre-training approach to precondition and identify entities without annotated datasets, then applies transfer learning models to different simulated limited datasets for a named entity recognition task.
Abstract
This paper explores the challenges and the methodologies associated with learning quality representations in scenarios with unlabelled small or limited datasets for downstream information extraction task (Multidomain Named Entity Recognition (NER). The study adopts a Transfer Learning on small datasets. Traditional NER systems often rely on large, labelled data, which is impractical for many domains. This study, therefore, applies an unsupervised pre-training approach to precondition and identify entities without annotated datasets, then applies transfer learning models to different simulated limited datasets for a named entity recognition task. Entity Recognition (NER) is essential in natural language processing (NLP), it identifies and classifies related entities within the text. This study addresses the complexities of domain variability, data sparsity, and overfitting and investigates innovative approaches such as data augmentation, few-shot learning, and domain adversarial training. Integrating these techniques promises to enhance the performance and generalizability of NER systems across diverse and resource-constrained domains, paving the way for more efficient and adaptable NLP applications.
This paper proposes to use mBERT IndicBERT and XLM-RoBERTa language models pretrained on target and other related Indic language corpora and further fine-tune these models for the NER task.
P. M, H. G, S. N· International journal of res...· 0 citations
The NerAxom dataset is presented, a BIO-tagged NER dataset for Assamese comprising 4,173 sentences and 106,046 tokens annotated across seven entity categories, and a set of language-specific post-processing rules based on morphological suffixes and keyword cues are introduced.
Punam Sarmah, M. Lahkar, Shobhanjana Kalita et al.· Engineering Research Express· 0 citations
INTRODUCTION
Clinical text in electronic health records contains valuable information for medical research. Named entity recognition (NER) and relation extraction (RE) are key tasks for structuring this information, but existing deep learning approaches are often computationally expensive.
METHODS
We develop a lightw...
Sihan Wu, Susanna Gordon, M. Boeker et al.· Studies in Health Technology...· 0 citations
The proposed AGPNer integrates a heterogeneous dependency fusion encoder, which reconstructs masked entities to enhance token representations and fine-tunes a hybrid dependency modeling block to learn domain-specific patterns in medical texts; an imbalance-adaptive span decoder, which decouples entity and non-entity sp...
Yuling Li, Yang-Juan Hu, Yi-Ming Bao et al.· ACM Transactions on Intellig...· 0 citations
Automatic named entity recognition (NER) is essential for many natural language processing applications, particularly in low-resource languages and those with corpora restricted to a single domain, where the lack of diverse data may hinder cross-domain generalisation. In this context, we attempt to contribute to the...
A. Sarymsakova, Ettore Mariotti, Helena Pérez Puente et al.· Natural Language Processing· 0 citations