This work introduces ConstructCIE, a manually annotated dataset for Causal Information Extraction from OSHA construction accident reports, and evaluates supervised sequence taggers and instruction-tuned LLMs in an end-to-end hierarchical extraction setting.
Abstract
Construction accident narratives contain rich causal information, but the evidence is often implicit, long-span, and distributed. We introduce ConstructCIE, a manually annotated dataset for Causal Information Extraction from OSHA construction accident reports. The dataset uses a hierarchical schema for accident types, causal factors, sub-causal factors, and supporting evidence spans. We evaluate supervised sequence taggers and instruction-tuned LLMs in an end-to-end hierarchical extraction setting. Results show that most evaluated models achieve strong accident-type prediction and recover broad causal meaning but remain limited in precise span-level extraction. Joint Hierarchical Extraction generally achieves stronger exact and soft matching, while Individual Hierarchical Extraction sometimes achieves higher keyword F1. Error distributions vary by extraction strategy, but evidence-selection and span-boundary errors remain common. These findings show that reliable Causal Information Extraction for construction accidents requires stronger domain grounding and more accurate evidence extraction. The code and data can be found at https://github.com/lab-flair/ConstructCIE .
Understanding the causes of phenomena is a key goal of scientific enquiry. For historic reasons, social scientists are reluctant to explicitly discuss causal assumptions. Nevertheless, causal claims do appear in the literature. Systematically cataloging these causal claims and parsing them into structured cause-effect...
Rasoul Norouzi, Bennett Kleinberg, Jeroen K. Vermunt et al.· Social science computer revi...· 0 citations
The Temporal Coherence Score (TCS), a continuous, intrinsically interpretable metric that quantifies the temporal coherence of a news article, is introduced by a four-stage pipeline: extraction of temporal facts, construction of a temporal knowledge graph, hierarchical verification against internal consistency rules an...
U.S. employment-discrimination complaints describe complex event sequences that are not explicitly captured by lexical or embedding-based representations alone. We present ARGUS, a source-grounded pipeline that combines a 5W1H-inspired schema, legal-domain models, and LLM-based structured generation to construct docume...
Structured extraction from Chinese military news supports intelligence analysis, decision-making, and knowledge base construction. However, existing resources provide limited support for joint informa?tion extraction in this domain, especially when events, event arguments, entities, and relations must be modeled togeth...
Yan-Cen Yu, Meng-Na Zhu, Zhen-Yu Song et al.· 0 citations
: Event knowledge graphs support event-centric question answering by linking events to temporal, location, participant, and source-record information, but semantic relevance alone does not guarantee that a selected event satisfies every represented condition. We propose EviGraphRAG, a ranker-agnostic reliability layer...
Yu-Teng Sun, Yang Su, Xu-An Wang· Computers, Materials & C...· 0 citations
This work proposes a guideline-based, three-stage LLM annotation framework for event extraction that incorporates detailed event annotation guidelines and supports multiple LLM annotators to improve robustness, and demonstrates that augmenting existing datasets with LLM-generated argument annotations can improve argume...
Marcel Geromel, Philipp Cimiano· Frontiers in Artificial Inte...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.