Skip to content

LLM-Structured Business Risk Disclosures from Japanese Annual Securities Reports for Long-Horizon Equity Risk Forecasting

Sep 2026 · IEEE Conference on Computational Intelligence for Financial Engineering & Economics · pp. 258-267 · 0 citations · 38 references

Abstract

This study evaluates whether large language models (LLMs) can convert the "Business and other risks" sections of Japanese annual securities reports into interpretable signals for long-horizon equity risk forecasting. Each extraction returns schema-constrained JSON, which is flattened into ordinal, binary, category-level, and temporal tabular features. For TOPIX 500 constituents, annual walk-forward tests compare price histories, fold-fitted character 3–5-gram TF–IDF, GPT-5.4-mini- and GPT-5.4-structured features, and their combinations. Price-history Ridge remains strongest for future volatility and for the annualized root-mean-square of negative daily returns. The clearest LLM result concerns maximum drawdown magnitude: temporal GPT-5.4 features improve a fixed nonlinear price model and add information beyond the current-year TF–IDF block, although the small number of annual folds limits year-level inference. Cross-fitted residual models show positive associations between disclosure features and price-model errors, but training-only shrinkage assigns little weight to LLM-only corrections; TF–IDF provides the most reliable final correction. Distributional diagnostics also reveal limited usable variation across several LLM scales: seven of twelve ordinal fields place at least 60% of documents on a single score. The evidence positions LLM structuring as an interpretable, target- and model-dependent complement to price and lexical representations rather than a universally superior forecasting representation.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.