Skip to content
Preprint

FinVerse: Financial Time-Series Benchmark

Aug 2026 · 1 citation · 24 references
Computer Science

TL;DR

FinVerse is introduced, a finance-domain time-series forecasting benchmark that takes a first step toward more realistic evaluation and highlights the need for domain-aware benchmarks that evaluate models under objectives closer to real-world decision making.

Abstract

As time-series foundation models have emerged, the need for benchmarks that can evaluate their forecasting ability in meaningful ways has become increasingly important. Existing time-series forecasting benchmarks provide useful standardized comparisons, but they often evaluate heterogeneous series with uniform error-based metrics. Strong performance under such metrics does not necessarily imply that a model's forecasts will support the best real-world decisions across domains. For example, in stock forecasting, correctly predicting whether a price will rise or fall can be more directly relevant to realized returns than minimizing point-wise forecast error alone. To this end, we introduce FinVerse, a finance-domain time-series forecasting benchmark that takes a first step toward more realistic evaluation. The released FinVerse data artifact contains 116,897 financial time series with 171.1M observations, of which 60,232 series with 17.4M observations are selected as evaluated targets based on their economic relevance to financial decisions. Unlike generic forecasting benchmarks that primarily emphasize uniform point-forecast or probabilistic accuracy, FinVerse defines 11 metric families comprising 78 evaluation metrics and assigns the most appropriate evaluation metrics to each individual time series based on its underlying economic meaning. Our analysis of 43 public time-series forecasting foundation models shows that strong performance under generic forecasting criteria does not necessarily translate into useful financial forecasts. This finding highlights the need for domain-aware benchmarks that evaluate models under objectives closer to real-world decision making.

View source

Similar papers

Case report Aug 2026

An Open Benchmark for Evaluating Time Series Forecasting Methods Across Financial Markets

Holding the data fixed and evaluating roughly a dozen univariate methods without exogenous regressors, it is found that asset returns remain near unforecastable across every method family and that hybrid and machine learning methods exhibit additional forecasting power on basis spreads and bank indicators.

J. Bejarano, Viren Desai, K. Keshava et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Beyond MSE: Rethinking the Evaluation Metric and Benchmarking for Irregular Time Series Forecasting

The Continuous-time Squared Error (CSE) is proposed, which employs importance weighting to eliminate the influence of the timestamp sampling distributions and theoretically proves that CSE's asymptotic estimation error with respect to continuous-time risk is no greater than that of MSE.

Rong Li, Haixin Xie, Xiao Wang et al. · 0 citations
Preprint Aug 2026

Long-Horizon Forecasting of Complete Financial Statements with Forma

Specialist training beats generalist scale when forecasting financial statements. To our knowledge, no prior work jointly forecasts complete financial statements beyond one year, yet in a discounted-cash-flow valuation most firm value sits past that window. We release ProForma-20Q, a reproducible benchmark for forecasting 78 statement line items 1-20 quarters ahead, for anonymized firms, from past statements and an industry code, scored by change-space $R^2$. On it, Forma, a transformer that reads statements as sets of (account, quarter, value) tuples and maximizes a masked-tuple Gaussian likelihood, beats every competitor we field: classical machine learning, chained gradient boosting, a zero-shot time-series foundation model, and frontier large language models. Its lead widens with horizon, where valuation needs accuracy most, and its Gaussian predictive intervals never under-cover. Forma's forecasts nearly satisfy accounting identities; exact coherence is recoverable at no statistically significant accuracy cost. Its tuple interface supports scenario analysis without retraining, and we show that pinning future revenue paths sharpens the rest of the statement.

Travis L. Johnson, Jian Jiang, Soumyabrata Chaudhuri et al. · 0 citations
Preprint Aug 2026

Forecast Collapse in Time-Series Foundation Models

When forecasting hourly returns for 1,000 US equities, we observe an unexpected phenomenon: predictions become nearly flat and show poor stock ranking, as measured by cross-sectional correlation. We call this forecast collapse. Surprisingly, the phenomenon largely disappears when forecasting trading volume under the same setting. We investigate forecast collapse across time-series foundation models (TSFMs), twelve deep-learning forecasting models, and 97 public benchmark configurations, and find that it is closely tied to target predictability. We identify two distinct reasons behind it: low predictability limits the amplitude of calibrated point forecasts, while per-series objectives leave cross-series structure unidentified. These findings reveal a calibration-ranking tradeoff: optimizing squared error leads to flat predictions, whereas directly optimizing cross-sectional correlation improves ranking but can inflate forecast amplitude by more than an order of magnitude. To address this tradeoff, we introduce CalibRank, a simple objective that balances calibration and ranking. On Finance1K, CalibRank nearly triples cross-sectional correlation while keeping amplitude close to the target, and improves correlation on all tested models. Our results reveal a blind spot in conventional time-series evaluation: per-series metrics can hide failures in cross-series structure needed by downstream decisions.

Shu Wan, Miles Ma, Hankun Zhu et al. · 0 citations
Preprint Sep 2026

Does Training on Future Data Pay? Look-Ahead Bias in Forecasting with Pretrained Models

We examine whether post-origin training information inflates the measured accuracy and economic value of financial forecasts. We evaluate five sets of financial time-series foundation models, each comprising independently trained annual vintages under U.S., global, and factor-augmented training environments, across 14 equity markets and four forecast horizons. Rolling comparisons vary the annual vintage for a fixed forecast; fixed-vintage comparisons hold the vintage fixed as target windows move across its training cutoff. Each alternative forecast is paired with an origin-aligned point-in-time (PIT) benchmark using identical numerical histories and inference protocols. In the U.S.-trained reference environment, post-origin vintages materially revise informative PIT forecasts but generally reduce accuracy in both designs. Pooled rolling comparisons yield higher mean squared forecast errors in 18 of 20 U.S. model-set-horizon combinations. The origin-crossing update also performs worse on average than an equally long pre-origin update. Under a common constrained allocation rule using one-month forecasts, median exposed-minus-PIT differences in annualized certainty-equivalent returns are -1.77 percentage points in the United States and -2.14 points internationally. Global and factor-augmented training produce more mixed predictive effects. An exact squared-error decomposition shows that revisions improve accuracy when their error-correcting benefit exceeds their mean squared magnitude; under U.S. training, alignment with PIT errors generally falls short of this requirement. Temporal exposure therefore establishes an information-set violation, not sufficient evidence of inflated predictive accuracy or investor value.

Hai-Qiang Chen, Li Chen, Yun-Long Chen et al. · 0 citations
Preprint Aug 2026

Multivariate Time Series Forecasting needs Cross Variable Loss

This work proposes CvLoss, a plug-in structural regularizer that constrains forecast residuals on a cross-variable graph and shows that CvLoss consistently improves competitive forecasting models, outperforms representative learning objectives, and is compatible with a variety of forecasting backbones.

Kuiye Ding, Yifan Hu, Hanchen Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.