Skip to content
Open access

Improving LLM-based event extraction with annotation guidelines

Aug 2026 · Frontiers in Artificial Intelligence · Vol 9 · 0 citations · 83 references
Medicine

TL;DR

This work proposes a guideline-based, three-stage LLM annotation framework for event extraction that incorporates detailed event annotation guidelines and supports multiple LLM annotators to improve robustness, and demonstrates that augmenting existing datasets with LLM-generated argument annotations can improve argument extraction performance under soft-matching evaluation.

Abstract

Event extraction constitutes a foundational task in information extraction, but reliance on laborious and expensive human annotations severely restricts the availability of training datasets. While recent works have explored Large Language Models (LLMs) as example-driven (or zero-shot) annotators, they are substantially outperformed by supervised techniques on structured extraction tasks, such as event detection and argument extraction, possibly on account of underspecified task instructions. In this work, we investigate to which extent LLMs can benefit from dataset-specific, detailed annotation guidelines that more precisely represent the dataset's underlying (human) annotation procedures. To this end, we propose a guideline-based, three-stage LLM annotation framework for event extraction that incorporates detailed event annotation guidelines and supports multiple LLM annotators to improve robustness. Using the comprehensive and well-documented ACE 2005 English Annotation Guidelines for Events as a reference document, we evaluate four LLMs across three guideline-compliant benchmark datasets. Our findings indicate that, depending on model choice, guideline specificity, and the dataset's relative label accuracy, employing detailed guidelines can considerably boost event extraction performance, gaining up to 6.7 F1 points over commonly used bare-minimum instructions, with particularly remarkable improvements for reasoning-based models. Furthermore, we demonstrate that augmenting existing datasets with LLM-generated argument annotations can improve argument extraction performance under soft-matching evaluation. Overall, our experiments emphasize the importance of annotation guidelines (as well as their specificity) for LLM-based annotations, providing valuable insights on leveraging LLMs as guideline-compliant annotators.

Read PDF

Similar papers

Open access 2026

Legal NER: Evaluating the Impact of LLM-Generated Annotations on NER Performance for Administrative Decisions

This study investigates the use of LLM-generated annotations to expand the training set for supervised NER models applied to sentences from Dutch administrative decisions as a low-resource domain and language and indicates that LLMs can accurately generate annotations for legal entities that are explicitly defined in l...

H. Nan, Samaneh Khoshrou, Johan Wolswinkel · 0 citations
Conference Sep 2026

LLM-based automatic identification and early warning of construction safety risks

This work proposes a hybrid framework that integrates prompt-guided attention with lightweight supervised fine-tuning to extract structured risk triples from heterogeneous construction texts and contributes to safer engineering practices through advanced NLP techniques.

Jin-Fei Liu, Qun Luo, Dou-Dou Li et al. · 0 citations
Open access Sep 2026

Using Large Language Models for Automated Corpus Annotation and Linguistic Analysis: A Critical Methodological Framework

Large language models (LLMs) are increasingly used to classify, label, summarize, and interpret large text collections, creating new possibilities for corpus linguistics. Their capacity for zero-shot and few-shot instruction following could reduce the cost of linguistic annotation and extend analysis beyond the categor...

Maria Ibrar · 0 citations
Preprint Sep 2026

Object Detection Benchmarks are Incomplete: The Role of Label Errors and Annotation Uncertainty

This work introduces a scalable annotation pipeline that emphasizes high recall and captures ambiguity through soft labels aggregated from at least 11 annotators per object, and shows that benchmark performance is highly sensitive to annotation quality, although model rankings remain largely stable.

Sarina Penquitt, Jonathan Klees, Antonia van Betteray et al. · 0 citations

Improving large-scale DLA datasets through semantic validation and relation-aware multimodal LLMs

A preliminary version of a framework that improves the most widespread DLA datasets quality by assessing and correcting layout coherence in scholarly documents and provides a more reliable ground truth with improved structural and semantic coherence for training and evaluating document segmentation and understanding mo...

L. Massai, S. Marinai · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.