Skip to content
Open access

Legal NER: Evaluating the Impact of LLM-Generated Annotations on NER Performance for Administrative Decisions

2026 · Journal of Computational Law and Legal Technology · 0 citations

TL;DR

This study investigates the use of LLM-generated annotations to expand the training set for supervised NER models applied to sentences from Dutch administrative decisions as a low-resource domain and language and indicates that LLMs can accurately generate annotations for legal entities that are explicitly defined in legislation but generate less reliable annotations for other legal entities that require deeper contextual understanding.

Abstract

Named entity recognition (NER) is a core information extraction (IE) task dependent on high-quality annotated data that are expensive and time-intensive to produce. Large language models (LLMs) offer a promising alternative through LLM-generated pseudo-annotations, yet their reliability in domain-specific legal settings remains insufficiently studied. This study investigates the use of LLM-generated annotations to expand the training set for supervised NER models applied to sentences from Dutch administrative decisions as a low-resource domain and language. To this end, LLM-based annotations of predefined legal entities are created using a schema-driven few-shot prompt, which are evaluated against a human-annotated dataset. Next, two different NER architectures are trained—a token-level NER model and a span-based NER–RE model (joint NER and relationship extraction (RE))—under three training settings: (1) human annotations only, (2) LLM-generated annotations only, and (3) models trained on LLM-generated annotations further fine-tuned on human annotations. The results indicate that LLMs can accurately generate annotations for legal entities that are explicitly defined in legislation but generate less reliable annotations for other legal entities that require deeper contextual understanding beyond what is explicitly stated in the text or prompt. Furthermore, fine-tuning a NER model trained on these LLM-generated annotations with human annotations turns out to slightly outperform models trained on human-annotated data only. Our findings highlight the potential of hybrid supervision strategies to scale low-resource legal NER tasks while maintaining human-level accuracy.

Read PDF

Similar papers

Open access Aug 2026

NerAxom: a BIO-tagged NER dataset and hybrid neural-rule framework for Assamese

The NerAxom dataset is presented, a BIO-tagged NER dataset for Assamese comprising 4,173 sentences and 106,046 tokens annotated across seven entity categories, and a set of language-specific post-processing rules based on morphological suffixes and keyword cues are introduced.

Punam Sarmah, M. Lahkar, Shobhanjana Kalita et al. · 0 citations
Conference Aug 2026

Automated Argument Mining and Entity Recognition in Sri Lankan Legal Judgments Using Large Language Models

Access to legal information remains a significant challenge in low-resource jurisdictions, where judicial documents are often lengthy, unstructured, and difficult to interpret. This paper presents a unified framework for automated argument mining and named entity recognition in Sri Lankan legal judgments using Large La...

Sathmi Sansiluni Jayaratne, Ruvan Weerasinghe · 0 citations
Preprint Aug 2026

ANNOTARES: A Dataset for Extracting Logical Structures from German Statutory Texts

This paper introduces the task of identifying and segmenting legal conditions (Tatbestand) and legal consequences (Rechtsfolge) within German statutory texts and presents ANNOTARES (Annotations of Tatbestand-Rechtsfolge Sequences), a novel dataset comprising German law texts with span-level annotations.

R. Schwarz, Jannik Strötgen · 0 citations
Open access Aug 2026

Biomedical Text Mining and Information Extraction Using Prompt-Enhanced and LoRA-Adapted Large Language Models

Biomedical named entity recognition (NER) and relation extraction (RE) remain challenging because biomedical texts contain ambiguous abbreviations, complex entity boundaries, domain-specific terminology, and implicit relations. This study proposes a prompt-enhanced and QLoRA-adapted large language model framework for b...

Feng Yan, De-Quan Zheng, Feng Yu et al. · 0 citations
#artificial intelligence Review Sep 2026

IntLawNER: A Named Entity Recognition Dataset and Benchmark in International Law

International law provides the normative framework through which states coordinate action, regulate armed conflict, and protect human rights, yet its texts remain without token-level named entity recognition (NER) resources. We introduce IntLawNER, a NER dataset and benchmark for codified sources of international law,...

Genis Skura, Roland Bouffanais, D. Wernli · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.