FWBench enables reproducible evaluation of how language models select and use time-series forecasts to make decisions under cost constraints.
Abstract
Time-series foundation models (TSFMs) provide forecasts for operational decisions, but accuracy alone does not determine their value. Evaluating agents that use these models requires measuring decision quality and forecast cost. FWBench evaluates this capability on 1,251 electricity and cycle-hire cases using fixed forecast tools and simulated capacity contracts. Agents select models, histories and horizons, then submit capacities to minimize a stated loss-cost objective. We evaluated two hosted and eight local configurations, including small language models, and tested local models with and without TSFMs. GPT-6 Astra bought inexpensive short-horizon forecasts selectively, using 2.5% of the budget, and outperformed fixed policies when the saved decisions were scored with three loss-cost weightings. FWBench enables reproducible evaluation of how language models select and use time-series forecasts to make decisions under cost constraints.
Operational forecasts are frequently updated as new data, assumptions, scenarios, and model configurations become available. In decision-facing forecasting workflows, the practical question is not only whether a forecast has changed, but whether the change is stable, supported by sufficient confidence, and actionable....
Rijul Saini, Gaurav Dewan, Shivani Sharma et al.· 0 citations
The results show that data-driven parameter optimization can guide LLM-based search over a broad space of inventory policy classes and identify high-performing, interpretable, and transferable decision rules.
Feng Yang, Preet Baxi, Yi Zhang et al.· 0 citations
FinRiskAtlas is introduced, a Chinese-language benchmark that evaluates financial LLMs along two complementary dimensions: operation execution under fixed evidence states and evidence-state control under evolving review conditions, and shows that broad financial capability scores do not fully capture where models are r...
Su-Yang Zhong, Jingzhe Zhu, Qi Xu et al.· 1 citation
We introduce Forecast-Dojo, a replayable environment for benchmarking and training LLM forecasting agents. It combines resolved prediction-market questions with dated news, allowing agents to research an event and revisit their predictions at successive historical dates. The same tasks and tools support repeated evalua...
Li-Qin Ye, Hao-Rui Wang, Fardin Ahmed et al.· 0 citations
General decision models, such as Jev, have recently emerged as efficient alternatives to LLMs for structured judgment and selection. But what kinds of decisions can these models reliably make, and how does their behavior change when individual decisions are composed into larger systems? To study this, we introduce JEVa...
Fei-Yu Duan, Jia-Yu Lin, Jia Wang et al.· 1 citation
CAS (causal active sequential experimentation), which targets evaluation to model-workload pairs and repeats the test as evidence accumulates, to ask whether one assignment stays optimal across every quality table consistent with the evidence.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.