Skip to content
Open access

PAP_NER: A large-scale vietnamese administrative named entity recognition corpus and hybrid deep learning architecture

Jul 2026 · PLoS ONE · Vol 21 · 0 citations · 43 references
Medicine

TL;DR

PAP_NER is presented, the first large-scale, gold-standard Vietnamese administrative NER corpus comprising 162,801 sentences with 205,807 entity annotations across five entity types critical for e-Government workflows, with implications for low-resource language NLP research.

Abstract

Named Entity Recognition (NER) is fundamental for automating administrative document processing in digital government systems. However, Vietnamese NLP research faces a critical infrastructure gap: existing datasets focus on generic information extraction (news, medical) rather than domain-specific administrative text. We present PAP_NER, the first large-scale, gold-standard Vietnamese administrative NER corpus comprising 162,801 sentences with 205,807 entity annotations across five entity types critical for e-Government workflows: Agency (CQ), Legal Document (VBPL), Object (ĐT), Datetime (NG), and Quantity (SL). The dataset was constructed through a rigorous human-in-the-loop annotation pipeline, achieving an inter-annotator agreement of κ = 0.85. We demonstrate PAP_NER’s value through comprehensive benchmarking of an established hybrid deep learning architecture, PhoBERT-CRF, which couples monolingual Transformer embeddings (PhoBERT) with Conditional Random Fields for structured prediction. PhoBERT-CRF achieves 97.95% Micro F1-score on the PAP_NER test set, significantly outperforming established baselines: BiLSTM+CRF (+2.01%), multilingual XLM-RoBERTa (+2.52%), and pure Transformer approaches (+0.44%). Ablation analysis reveals that the CRF layer provides statistically significant improvements for structurally complex entities (VBPL: + 0.96%, p < 0.05, McNemar’s test). We release PAP_NER publicly (DOI: 10.5281/zenodo.18044019) under Creative Commons BY 4.0 license to support reproducibility and enable further research in Vietnamese administrative NLP. This work establishes a foundational dataset and methodology for addressing the Vietnamese government NER gap, with implications for low-resource language NLP research.

Read PDF

Similar papers

Open access 2026

Legal NER: Evaluating the Impact of LLM-Generated Annotations on NER Performance for Administrative Decisions

This study investigates the use of LLM-generated annotations to expand the training set for supervised NER models applied to sentences from Dutch administrative decisions as a low-resource domain and language and indicates that LLMs can accurately generate annotations for legal entities that are explicitly defined in l...

H. Nan, Samaneh Khoshrou, Johan Wolswinkel · 0 citations
Open access Aug 2026

Sequence Labeling in Urdu Social Media Texts: Data Annotation and Transformer-Based Deep Learning Models

A transformer-based hybrid architecture is proposed, XLM-R–CNN–BiLSTM–CRF, which integrates XLM-R embeddings for contextualized representation, convolutional neural networks (CNNs) for local feature extraction, bidirectional long short-term memory (BiLSTM) networks for sequential modeling, and a CRF layer for optimal s...

Rafiul Haq, Xiao-Wang Zhang, Sofonias Yitagesu et al. · 0 citations
Open access 2026

Kannada Named Entity Recognition Using Deep Learning Techniques

This paper proposes to use mBERT IndicBERT and XLM-RoBERTa language models pretrained on target and other related Indic language corpora and further fine-tune these models for the NER task.

P. M, H. G, S. N · 0 citations
Open access Aug 2026

NerAxom: a BIO-tagged NER dataset and hybrid neural-rule framework for Assamese

The NerAxom dataset is presented, a BIO-tagged NER dataset for Assamese comprising 4,173 sentences and 106,046 tokens annotated across seven entity categories, and a set of language-specific post-processing rules based on morphological suffixes and keyword cues are introduced.

Punam Sarmah, M. Lahkar, Shobhanjana Kalita et al. · 0 citations
#natural language process... Preprint Aug 2026

En-ViMedNER: An English-Vietnamese Parallel Biomedical Corpus with UMLS Semantic Type Annotations

En-ViMedNER is presented, the first English-Vietnamese parallel biomedical NER corpus annotated with UMLS semantic types, which are language-neutral codes providing a shared cross-lingual label space and ensuring direct comparability with existing UMLS-based resources.

Nhu Vo, P. Nguyen, Nu-Uyen-Phuong Le et al. · 0 citations
Open access Aug 2026

AraCTI-NER: A Dataset and Benchmark for Arabic Cyber Threat Intelligence Named Entity Recognition

A AraCTI-NER is introduced, a dataset of 10,312 token-level annotated samples over eight STIX-inspired entity types built by an LLM-assisted pipeline seeded with authentic Arabic cybersecurity articles, structurally validated and rebalanced through targeted generation.

Joud Alghamdi, Souham Meshoul · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.