Skip to content
Preprint

ConstructCIE: A Dataset for Extracting Causal Information from Construction Accident Narratives

Aug 2026 · 0 citations · 42 references
Computer Science

TL;DR

This work introduces ConstructCIE, a manually annotated dataset for Causal Information Extraction from OSHA construction accident reports, and evaluates supervised sequence taggers and instruction-tuned LLMs in an end-to-end hierarchical extraction setting.

Abstract

Construction accident narratives contain rich causal information, but the evidence is often implicit, long-span, and distributed. We introduce ConstructCIE, a manually annotated dataset for Causal Information Extraction from OSHA construction accident reports. The dataset uses a hierarchical schema for accident types, causal factors, sub-causal factors, and supporting evidence spans. We evaluate supervised sequence taggers and instruction-tuned LLMs in an end-to-end hierarchical extraction setting. Results show that most evaluated models achieve strong accident-type prediction and recover broad causal meaning but remain limited in precise span-level extraction. Joint Hierarchical Extraction generally achieves stronger exact and soft matching, while Individual Hierarchical Extraction sometimes achieves higher keyword F1. Error distributions vary by extraction strategy, but evidence-selection and span-boundary errors remain common. These findings show that reliable Causal Information Extraction for construction accidents requires stronger domain grounding and more accurate evidence extraction. The code and data can be found at https://github.com/lab-flair/ConstructCIE .

View source

Similar papers

Open access Sep 2026

Parsing Causal Relationships in Social Science Publications

Understanding the causes of phenomena is a key goal of scientific enquiry. For historic reasons, social scientists are reluctant to explicitly discuss causal assumptions. Nevertheless, causal claims do appear in the literature. Systematically cataloging these causal claims and parsing them into structured cause-effect...

Rasoul Norouzi, Bennett Kleinberg, Jeroen K. Vermunt et al. · 0 citations
#artificial intelligence Preprint Aug 2026

An Explainable Coherence Score for Detecting Temporal Inconsistencies in Political News

The Temporal Coherence Score (TCS), a continuous, intrinsically interpretable metric that quantifies the temporal coherence of a news article, is introduced by a four-stage pipeline: extraction of temporal facts, construction of a temporal knowledge graph, hierarchical verification against internal consistency rules an...

M. Pantea, Adrian Groza · 0 citations
#natural language process... Preprint Sep 2026

ARGUS: Role-Aware Event Knowledge Graphs for U.S. Employment-Discrimination Complaints

U.S. employment-discrimination complaints describe complex event sequences that are not explicitly captured by lexical or embedding-based representations alone. We present ARGUS, a source-grounded pipeline that combines a 5W1H-inspired schema, legal-domain models, and LLM-based structured generation to construct docume...

Sriram Kannan, Swetha Saseendran, Vishnu Vardhan Reddy Kandi et al. · 0 citations
#natural language process... Preprint Sep 2026

CMNIE: An Information Extraction Benchmark for Chinese Military News

Structured extraction from Chinese military news supports intelligence analysis, decision-making, and knowledge base construction. However, existing resources provide limited support for joint informa?tion extraction in this domain, especially when events, event arguments, entities, and relations must be modeled togeth...

Yan-Cen Yu, Meng-Na Zhu, Zhen-Yu Song et al. · 0 citations
Open access 2026

EviGraphRAG: Constraint-Verified Retrieval-Augmented Question Answering over Event Knowledge Graphs

: Event knowledge graphs support event-centric question answering by linking events to temporal, location, participant, and source-record information, but semantic relevance alone does not guarantee that a selected event satisfies every represented condition. We propose EviGraphRAG, a ranker-agnostic reliability layer...

Yu-Teng Sun, Yang Su, Xu-An Wang · 0 citations
Open access Aug 2026

Improving LLM-based event extraction with annotation guidelines

This work proposes a guideline-based, three-stage LLM annotation framework for event extraction that incorporates detailed event annotation guidelines and supports multiple LLM annotators to improve robustness, and demonstrates that augmenting existing datasets with LLM-generated argument annotations can improve argume...

Marcel Geromel, Philipp Cimiano · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.