Skip to content
Open access

From incident narratives to actionable controls: insights on the iron & steel industry using LLM assisted learning from incident databases

2026 · AHFE International · 0 citations

Abstract

Learning from incidents is a cornerstone of occupational safety risk management, especially in high-risk industrial sectors. Incident databases support this process by collecting records that describe the causes, dynamics, and consequences of adverse events. However, these databases largely rely on unstructured textual narratives, which limits systematic analysis and the translation of learned lessons into effective preventive actions. This paper focuses on the iron and steel industry and presents an analysis pipeline supported by Large Language Models (LLMs) for extracting, synthesising, and structuring information from two major incident databases: the U.S. OSHA database and the French ARIA database. Relevant records were selected using industry classification codes and pre-processed to harmonise terminology, normalise information fields, remove duplicates, and manage multilingual content. Within a human-in-the-loop framework, LLMs were used to identify critical occupational risk scenarios, characterise them in terms of frequency and severity, and derive prevention and risk mitigation measures structured according to the ISO 45001 hierarchy of controls. Eight critical scenarios were identified and subsequently validated and refined by safety experts from the steel industry. Quantitative analysis identified point-of-operation machinery and load handling as the most frequent scenarios, while confined spaces and high-energy events exhibited disproportionate severity and lethality. The results demonstrated how LLM-supported approaches can enhance learning from incidents by transforming large volumes of heterogeneous narrative data into a traceable, expert-validated knowledge base that supports hazard identification, risk assessment and management, and continuous improvement in high-risk industrial environments.

Read PDF

Similar papers

Aug 2026

What is Missing? An HFACS Analysis of the VERIS Community Database

Human action drives most cybersecurity breaches, yet industry reports rely on vague labels that lack diagnostic utility. This study evaluates whether public incident narratives from the VERIS Community Database provide the context required for systemic intervention to 45 Action.Error records from 2020 to 2021 were analyzed across three error varieties: misdelivery, misconfiguration, and publishing. Using the Human Factors Analysis and Classification System, we applied a strict evidence-based coding strategy to map narrative to systemic levels. Our results reveal a significant diagnostic gap: while narratives consistently support Level 1 (Unsafe Act) coding, evidence for latent preconditions, supervision failures, and organizational influences remains largely absent. Specifically, misdelivery and misconfiguration align cleanly with skill-based and decision errors, but the causal chain stops at the individual. These findings suggest that current public reporting can identify the type of error that occurred, but is limited in explaining why it occurred.

Saroja Roy Grandhi, Jeremiah D. Still · 0 citations
Preprint Aug 2026

AISA: AI Safety Assistant Framework for Continuous Improvement of Highway Construction

Job Safety Analysis (JSA) and pre-task planning can benefit from prior incident records, yet historical accident data is often stored as unstructured narratives that are difficult to consult at the point of planning. A novel framework centered on large language models (LLMs) for highway construction safety reporting and planning is proposed as a foundation for future agentic applications, prioritizing deterministic, local inferencing. The first aim is to enable classification and quality scoring of incident narratives for existing and future reporting purposes. The second is to evaluate retrieval of relevant historical accidents, related imagery, and trusted industry documents for incorporation into daily safety plans. Neural probes were trained to classify incidents along four multiclass and two binary Occupational Injury and Illness Classification System (OIICS) fields and to derive an overall quality score, evaluated on a test set of over 15,000 narratives and a held-out set of 100 author-labeled records, benchmarked against a majority-vote LLM ensemble. The retrieval of historical accidents, reference imagery, and industry documents was benchmarked across embedding models using standard information retrieval metrics. OIICS classification reached 75% held-out accuracy, though the two binary flags were degenerate. The quality score, while meaningful on one database, was distorted on out-of-distribution fatalities in the held-out dataset. Accident retrieval recovered relevant incidents far above chance, performing best on lexically distinct construction activities. On document question answering, an open-weight decoder embedding model surpassed proprietary models. Overall, this work provides a new framework rooted in local inferencing and text embedding models for future agentic applications, with emphasis on bridging external data to JSA reports.

M. Smetana, Trevor Neece, Lev Khazanovich · 0 citations
Review Jul 2026

LAMDA: Large Language Model as Decision Analyst

Influence diagrams address the challenges of decision-making under risk by structuring information, decisions and values, while clearly depicting uncertainties and probabilistic dependencies. However, constructing an influence diagram requires expertise in decision analysis and is further complicated by the need to process large amounts of contextual information. This work focuses on the construction of influence diagrams from natural language input by leveraging large language models (LLMs). We design a workflow that prompts LLMs to output elements of an influence diagram and resolves issues through verification and regeneration. We also construct a new dataset of typical decision problems under risk. Evaluations using this dataset demonstrate that our framework effectively identifies key factors and relationships in natural language, making better decisions than standalone LLMs and LLMs enhanced with standard techniques such as chain-of-thought (CoT). Finally, we apply LAMDA to discussions by groups of disease control experts on a hypothetical pandemic to demonstrate its real-world applicability. Overall, the method effectively synthesizes unstructured text into an influence diagram that, while subject to human review and refinement, enhances information processing and supports decision-making.

Yifan Hong, Sizhong Qin, Chen Wang · 0 citations
Preprint Aug 2026

Evaluating Large Language Model Performance on International Maritime Dangerous Goods Code Compliance

The transport of dangerous goods by sea is a high-consequence activity governed by the International Maritime Dangerous Goods (IMDG) Code, a complex regulatory framework where errors in classification, packaging, stowage, or segregation can result in fire, explosion, toxic release, or loss of life or vessel. Correct compliance requires accurately interpreting hundreds of pages of interacting provisions, updated on a two-year amendment cycle. Practitioners increasingly use Large Language Models (LLMs) as decision-support tools, yet no systematic evaluation exists of whether they can reliably interpret IMDG requirements for safety-critical use. This paper introduces DGEval, the first benchmark for evaluating LLM knowledge of IMDG Amendment 42-24. Built from expert-written questions on the NCB Hazcheck e-learning platform and structured lookups from the Dangerous Goods List (DGL), it comprises 1,678 questions across multiple-choice, open-ended, DGL lookup, and regulatory identification tasks. We evaluate 13 models from six providers across multiple thinking configurations, including one maritime domain-specific fine-tuned model, and test the effect of web search. Although the best-performing model exceeds the human practitioner baseline on multiple-choice questions, all models are weakest in the operationally safety-critical areas of stowage, segregation, and regulatory recall. These results indicate that LLMs may support compliance tasks, particularly structured DGL lookups with web search, but unreliability in operational areas and regulatory-text recall means human oversight and authoritative source verification remain necessary before deployment in any safety-critical context. DGEval is designed as a safety assurance instrument to be applied continuously as models evolve, not as a settled characterisation of current capability.

Alexander S Thomas, Hubert P. H. Shum, Darren Nellis et al. · 0 citations
Review Open access Aug 2026

INTEGRATING MACHINE LEARNING WITH OCCUPATIONAL INCIDENT ANALYTICS: EMERGING FRAMEWORKS FOR NATIONWIDE WORKPLACE SAFETY ENHANCEMENT

The integration of predictive analytics into United States occupational safety and health practice promises to shift workplace risk management from post-incident recordkeeping toward prospective injury prevention. Synthesizing U.S.-focused research published between 2021 and 2026, this narrative review critically evaluates machine learning applications across severe incident classification, narrative text processing, real-time computer vision, return-to-work outcome forecasting, and emerging federal oversight models. While algorithmic capabilities have reached high computational performance, including production-scale transformer deployments for administrative coding and high-accuracy ensembles for accident narrative parsing, the literature remains dominated by retrospective offline experiments. Crucially, empirical evaluations demonstrate a profound gap between model precision and tangible worker safety, as virtually no published studies measure prospective reductions in workplace injury or illness rates. This lack of demonstrated field impact is further complicated by severe systemic data fragmentation across federal enforcement registries, statistical surveys, sector-specific databases, and state-bounded workers' compensation claims. To bridge this divide, this paper articulates an integrated nationwide framework featuring federated data interoperability, risk-calibrated algorithmic hygiene standards, mandatory prospective evaluation protocols, and a phased evolution toward binding administrative regulation. Aligning computational innovation with measurable workplace hazard reduction, rather than further optimizing classification accuracy on historical datasets, represents the essential mandate for the future of occupational safety analytics.

A. Alao · 0 citations