LLM-Structured Business Risk Disclosures from Japanese Annual Securities Reports for Long-Horizon Equity Risk Forecasting
Abstract
This study evaluates whether large language models (LLMs) can convert the "Business and other risks" sections of Japanese annual securities reports into interpretable signals for long-horizon equity risk forecasting. Each extraction returns schema-constrained JSON, which is flattened into ordinal, binary, category-level, and temporal tabular features. For TOPIX 500 constituents, annual walk-forward tests compare price histories, fold-fitted character 3–5-gram TF–IDF, GPT-5.4-mini- and GPT-5.4-structured features, and their combinations. Price-history Ridge remains strongest for future volatility and for the annualized root-mean-square of negative daily returns. The clearest LLM result concerns maximum drawdown magnitude: temporal GPT-5.4 features improve a fixed nonlinear price model and add information beyond the current-year TF–IDF block, although the small number of annual folds limits year-level inference. Cross-fitted residual models show positive associations between disclosure features and price-model errors, but training-only shrinkage assigns little weight to LLM-only corrections; TF–IDF provides the most reliable final correction. Distributional diagnostics also reveal limited usable variation across several LLM scales: seven of twelve ordinal fields place at least 60% of documents on a single score. The evidence positions LLM structuring as an interpretable, target- and model-dependent complement to price and lexical representations rather than a universally superior forecasting representation.