Holding the data fixed and evaluating roughly a dozen univariate methods without exogenous regressors, it is found that asset returns remain near unforecastable across every method family and that hybrid and machine learning methods exhibit additional forecasting power on basis spreads and bank indicators.
Abstract
Accurate time series forecasts underpin asset pricing, risk management, monetary and macroprudential policy, and other applications. The set of available forecasting methods is expanding rapidly, driven by new machine learning models. This raises a practical question: Do these methods deliver real forecasting power gains on financial data? Since no single method is best across all data generating processes, the question can be answered only by direct evaluation on domain-specific data, and a fair comparison requires holding the data fixed across methods. Yet, new methods are rarely evaluated on the financial data commonly used in academic finance; when they are, each is assessed on its own dataset and cleaning conventions. Apparent method rankings entangle method skill with cleaning choices that themselves require domain expertise. We address this by assembling, in one place, canonical cleaning procedures for many financial datasets to hold the data fixed across forecasting experiments. We introduce a standardized open-source dataset covering equities, corporate bonds, U.S. Treasuries, foreign exchange, commodities, credit default swaps, options, five basis spread datasets, and bank and intermediary indicators, each cleaned per the canonical paper for that asset class. Holding the data fixed and evaluating roughly a dozen univariate methods without exogenous regressors, we find that asset returns remain near unforecastable across every method family and that hybrid and machine learning methods exhibit additional forecasting power on basis spreads and bank indicators.
FinVerse is introduced, a finance-domain time-series forecasting benchmark that takes a first step toward more realistic evaluation and highlights the need for domain-aware benchmarks that evaluate models under objectives closer to real-world decision making.
Jaehoon Lee, Jun Seo, Seunghan Lee et al.· 1 citation
We examine whether post-origin training information inflates the measured accuracy and economic value of financial forecasts. We evaluate five sets of financial time-series foundation models, each comprising independently trained annual vintages under U.S., global, and factor-augmented training environments, across 14 equity markets and four forecast horizons. Rolling comparisons vary the annual vintage for a fixed forecast; fixed-vintage comparisons hold the vintage fixed as target windows move across its training cutoff. Each alternative forecast is paired with an origin-aligned point-in-time (PIT) benchmark using identical numerical histories and inference protocols. In the U.S.-trained reference environment, post-origin vintages materially revise informative PIT forecasts but generally reduce accuracy in both designs. Pooled rolling comparisons yield higher mean squared forecast errors in 18 of 20 U.S. model-set-horizon combinations. The origin-crossing update also performs worse on average than an equally long pre-origin update. Under a common constrained allocation rule using one-month forecasts, median exposed-minus-PIT differences in annualized certainty-equivalent returns are -1.77 percentage points in the United States and -2.14 points internationally. Global and factor-augmented training produce more mixed predictive effects. An exact squared-error decomposition shows that revisions improve accuracy when their error-correcting benefit exceeds their mean squared magnitude; under U.S. training, alignment with PIT errors generally falls short of this requirement. Temporal exposure therefore establishes an information-set violation, not sufficient evidence of inflated predictive accuracy or investor value.
Hai-Qiang Chen, Li Chen, Yun-Long Chen et al.· 0 citations
Research on stock market prediction relies on public datasets to develop and compare learning-based models. However, existing datasets are often limited to price-based information, cover a restricted set of assets or time periods, or provide only a subset of the heterogeneous signals required by modern prediction approaches. Moreover, differences in data preprocessing, task formulation, and evaluation protocols make experimental results difficult to reproduce and hard to compare across studies. To address these challenges, we present FinBench, a benchmarking framework designed to support reproducible evaluation of stock market prediction models using heterogeneous financial data. FinBench integrates historical prices, inter-company relations, macroeconomic indicators, and financial news within a unified data construction pipeline, and standardizes feature extraction and evaluation across classification, regression, and ranking tasks. Model outputs are evaluated both at the task prediction level and through portfolio-based strategies under consistent experimental settings. Using FinBench, we benchmark a range of existing models on European and US equity markets, analyzing their behavior across different prediction tasks, portfolio constructions, and investment universes with diverse liquidity and structural characteristics. FinBench provides a practical reference for systematic and reproducible comparison of stock market prediction methods.
Sara Pederzoli, Marta Santacroce, Francesco Guerra et al.· Proceedings of the 32nd ACM...· 0 citations
Specialist training beats generalist scale when forecasting financial statements. To our knowledge, no prior work jointly forecasts complete financial statements beyond one year, yet in a discounted-cash-flow valuation most firm value sits past that window. We release ProForma-20Q, a reproducible benchmark for forecasting 78 statement line items 1-20 quarters ahead, for anonymized firms, from past statements and an industry code, scored by change-space $R^2$. On it, Forma, a transformer that reads statements as sets of (account, quarter, value) tuples and maximizes a masked-tuple Gaussian likelihood, beats every competitor we field: classical machine learning, chained gradient boosting, a zero-shot time-series foundation model, and frontier large language models. Its lead widens with horizon, where valuation needs accuracy most, and its Gaussian predictive intervals never under-cover. Forma's forecasts nearly satisfy accounting identities; exact coherence is recoverable at no statistically significant accuracy cost. Its tuple interface supports scenario analysis without retraining, and we show that pinning future revenue paths sharpens the rest of the statement.
Travis L. Johnson, Jian Jiang, Soumyabrata Chaudhuri et al.· 0 citations
QFRS underpins a public state-of-the-art leaderboard, ensuring that only studies satisfying these standards are ranked, with the goal of shifting the literature from opaque, error-metric-driven results to transparent, economically meaningful and comparable benchmarks.
M. Khushi· Artificial Intelligence Revi...· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.