Skip to content

Forecast Workflow Bench: Evaluating Language-Model Decisions with Budgeted Forecast Tools

Sep 2026 · 0 citations · 23 references
Computer Science

TL;DR

FWBench enables reproducible evaluation of how language models select and use time-series forecasts to make decisions under cost constraints.

Abstract

Time-series foundation models (TSFMs) provide forecasts for operational decisions, but accuracy alone does not determine their value. Evaluating agents that use these models requires measuring decision quality and forecast cost. FWBench evaluates this capability on 1,251 electricity and cycle-hire cases using fixed forecast tools and simulated capacity contracts. Agents select models, histories and horizons, then submit capacities to minimize a stated loss-cost objective. We evaluated two hosted and eight local configurations, including small language models, and tested local models with and without TSFMs. GPT-6 Astra bought inexpensive short-horizon forecasts selectively, using 2.5% of the budget, and outperformed fixed policies when the saved decisions were scored with three loss-cost weightings. FWBench enables reproducible evaluation of how language models select and use time-series forecasts to make decisions under cost constraints.

View source

Similar papers

Book Open access Oct 2026

ForeACT: A Model-Driven Workbench for Actionability Assessment of Forecast Changes

Operational forecasts are frequently updated as new data, assumptions, scenarios, and model configurations become available. In decision-facing forecasting workflows, the practical question is not only whether a forecast has changed, but whether the change is stable, supported by sufficient confidence, and actionable....

Rijul Saini, Gaurav Dewan, Shivani Sharma et al. · 0 citations
Review Aug 2026

FinRiskAtlas: Decision-Aligned Evaluation of Large Language Models for Financial Risk Review

FinRiskAtlas is introduced, a Chinese-language benchmark that evaluates financial LLMs along two complementary dimensions: operation execution under fixed evidence states and evidence-state control under evolving review conditions, and shows that broad financial capability scores do not fully capture where models are r...

Su-Yang Zhong, Jingzhe Zhu, Qi Xu et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Forecast-Dojo: Replayable Environments for Benchmarking and Training LLM Forecasting Agents

We introduce Forecast-Dojo, a replayable environment for benchmarking and training LLM forecasting agents. It combines resolved prediction-market questions with dated news, allowing agents to research an event and revisit their predictions at successive historical dates. The same tasks and tools support repeated evalua...

Li-Qin Ye, Hao-Rui Wang, Fardin Ahmed et al. · 0 citations
#natural language process... Preprint Oct 2026

General Decision Models: Benchmarking and Insights Beyond Jev

General decision models, such as Jev, have recently emerged as efficient alternatives to LLMs for structured judgment and selection. But what kinds of decisions can these models reliably make, and how does their behavior change when individual decisions are composed into larger systems? To study this, we introduce JEVa...

Fei-Yu Duan, Jia-Yu Lin, Jia Wang et al. · 1 citation

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.