This study presents a novel, manually annotated dataset of 121 business news articles related to five major steel companies, using the Cambridge Risk Taxonomy, and provides empirical evidence on the opportunities and limitations of LLMs for both analytical and narrative forms of automated risk assessment.
Abstract
The increasing complexity of global e-commerce supply chains underscores the need for automated, context-aware risk monitoring systems capable of interpreting large volumes of unstructured information. Although Large Language Models (LLMs) have shown strong performance across natural language processing tasks, their application to real-world supply chain risk detection remains limited. This study presents a novel, manually annotated dataset of 121 business news articles related to five major steel companies, using the Cambridge Risk Taxonomy. Leveraging this dataset, we evaluate two state-of-the-art LLMs in a multi-label risk classification task using few-shot prompting. The results demonstrate that LLMs can approximate human annotation, though challenges persist in detecting domain-specific risks such as Geopolitical threats and in avoiding label overgeneration. Beyond classification, we further assess the capacity of LLMs to generate managerial risk summaries. We show that summaries derived from model-predicted risks exhibit strong semantic alignment to summaries generated from human annotations, highlighting the potential of LLMs to support executive-level risk interpretation. Overall, this study contributes the first publicly available dataset of fine-grained, hierarchical risk annotations in an e-commerce supply chain context and provides empirical evidence on the opportunities and limitations of LLMs for both analytical and narrative forms of automated risk assessment.
Empirical evaluation of TaSC-LLM shows that the proposed framework achieves strong taxonomy coverage, classification accuracy, and agreement with expert annotations, and suggests that LLM-based reasoning can help convert unstructured user-generated text into interpretable topic measures for downstream empirical and man...
Ensuring construction safety requires timely identification of latent risks embedded within unstructured documents such as inspection logs, incident reports, and supervisor notes. Traditional rule-based or statistical methods often struggle to extract such knowledge due to linguistic ambiguity, domain-specific expressi...
Jin-Fei Liu, Qun Luo, Dou-Dou Li et al.· International Conference on...· 0 citations
This work introduces RegNLI, a novel framework that formulates misbranding detection as a inference task between product claims and regulatory provisions, and builds a foundation for compliance-aware NLP systems and opens new directions for integrating formal reasoning with neural architectures in regulatory domains.
CorporateBench (CB), a human-validated multi-task Q&A benchmark whose scale approaches the conditions LLMs encounter in corporate communication networks, is presented, with evaluation corpora surpassing 230,000 documents.
Sil Hamilton, Albert Yu Sun, Oscar Romero et al.· 0 citations
This work introduces WildSEEK, a manually annotated dataset of 3k information-seeking queries from real user interactions, and an evaluation framework for LLM-generated responses, and finds that over a third of information-seeking queries are high-risk and more often analytical.
Tanise Ceron, Joachim Baumann, Elisa Bassignana et al.· 0 citations
A comprehensive natural language processing (NLP) pipeline for extracting key information and identifying industries, which outperforms traditional machine learning baselines and single deep learning models, offering more reliable recognition for minority classes.
Xin-Yi Xu· Applied and Computational En...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.