Jun 2026· arXiv.org· Vol abs/2607.18271· 0 citations· 68 references
Computer Science
TL;DR
Results show that generated explanations approached analyst-written explanations in terms of readability, consistency and persuasiveness, demonstrating that grounded explanation generation for time series forecasting can be achieved at scale without domain-specific fine-tuning.
Abstract
Time series forecasts are widely used in decision-critical domains, where they are rarely consumed without accompanying explanations. Producing such explanations is usually a manual and costly process, and attempts to automate it using large language models often suffer from hallucination when applied to temporal data. We propose a domain-agnostic framework for grounded natural language explanation generation for time series forecasts, illustrated in Figure 1. The framework consists of three components: (i) extraction of structured explanatory factors from historical analyst-written explanations, (ii) evidence-conditioned explanation generation, and (iii) scalable evaluation for readability, logical consistency, and persuasiveness. The design explicitly constrains generation to verifiable evidence, reducing unsupported claims. We evaluate the framework on a financial forecasting case study involving the NASDAQ-100 index and a freight pricing case study using data from Vortexa. Results show that generated explanations approached analyst-written explanations in terms of readability, consistency and persuasiveness. These findings demonstrate that grounded explanation generation for time series forecasting can be achieved at scale without domain-specific fine-tuning.
ReasonCast, the recipe for finetuning any LLM to perform both tasks jointly, yields a model that generates a reasoning chain and a forecast together in a single autoregressive pass, and Extensive experiments show that ReasonCast outperforms both LLMs and TS models on prediction accuracy while producing verifiable, caus...
Seung-Han Lee, Jun Seo, Jaehoon Lee et al.· 0 citations
Explaining deep learning models operating on time series data is crucial in various applications that require transparent and interpretable insights into model behavior. {Existing explanation methods generally fall into two categories: attribution-based explanations, which identify the temporal regions most responsible...
Xu Zheng, Zichuan Liu, Zhuomin Chen et al.· 0 citations
As forecasts increasingly drive decisions in fields such as energy, transportation, and healthcare, understanding the historical data behind these predictions has become as crucial as the predictions themselves. Although existing interpretable-by-design forecasters reveal their internal structures, they offer no guaran...
It is suggested that current LLMs often rely on heuristic arbitration strategies when integrating heterogeneous evidence, highlighting a failure mode for tool-augmented decision systems.
Mattia Carletti, Edward Phillips, Fredrik K. Gustafsson et al.· 0 citations
Inspired by Pearl's counterfactual notion of necessity, TimePNS assesses whether a temporal factor is necessary by intervening on it and measuring whether the original prediction is disrupted, and consistently improves sufficiency-necessity trade-offs over strong baselines.
Hongnan Ma, Yi-Wei Shi, Meng-Yue Yang et al.· arXiv.org· 0 citations
This work examines whether Large Language Models (LLMs) can serve as explanation layers that translate post-hoc explanation artefacts into stakeholder-appropriate risk narratives and discusses implications for the governance of risk models, including deployment considerations and the value of domain-aligned LLMs in reg...
Sahab Zandi, Noah Kostesku, Christophe Mues et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.