Skip to content
Preprint

Forecasting in the Fog: Real-Time versus Revised-Data Evidence on Machine Learning's Edge over the Phillips Curve

Aug 2026 · 0 citations · 6 references
Mathematics

TL;DR

Whether the ML advantage over the Phillips curve documented in Agyekum (2026) survives when models are trained and evaluated on real-time (ALFRED) vintages rather than revised series, and whether SHAP feature-importance rankings are an artifact of in-sample estimation.

Abstract

ML inflation forecasts are almost universally trained on fully revised data, even though real-time forecasters never have such data, and reported feature importances are typically computed in-sample, conflating predictive relevance with retrospective fit. This paper asks whether the ML advantage over the Phillips curve documented in Agyekum (2026) survives when models are trained and evaluated on real-time (ALFRED) vintages rather than revised series, and whether SHAP feature-importance rankings are an artifact of in-sample estimation. Using 2000-2026 U.S. data on unemployment, CPI and PCE inflation, payrolls, real GDP, and the 10-year-2-year Treasury spread, vintage-consistent panels are built for four traditional models (random walk, AR(1), Phillips curve, ADL-OLS) and four ML models (Random Forest, Gradient Boosting, Elastic Net, SVR), re-estimated recursively at 3-, 6-, and 12-month horizons (208, 206, 204 forecasts). Real-time/revised accuracy differences are small and, apart from one exception at 6 months (Gradient Boosting vs. Phillips curve, DM = -1.671, p = 0.097), indistinguishable under Diebold-Mariano tests; Gradient Boosting alone shows consistent positive skill at longer horizons. The random walk remains a strong short-horizon benchmark, consistent with the puzzle in Agyekum et al. (2026) for exchange rates. Using walk-forward, out-of-sample SHAP, a Random Forest on revised data assigns dominant importance to PCE inflation (mean |SHAP| = 0.778, rank 1 of 9), while on real-time data it assigns PCE negligible importance (0.039, rank 6), relying instead on current CPI (0.834 vs. 0.223). This twenty-fold swing, larger than the in-sample estimate, is invisible to point-forecast metrics and shows the model's PCE reliance is substantially a hindsight artifact. An RSI summarizes the accuracy gap by model and horizon, with implications for auditing ML inflation forecasts.

View source

Similar papers

Open access Sep 2026

Linear and Nonlinear Econometric Models versus Machine-Learning Models: Evidence from Realized-Volatility Forecasting

A robust horizon-dependent ranking is revealed: Markov-switching HAR performs best at short horizons, ARFIMA generally leads at the monthly horizon, and the five-day horizon is intermediate, and forecast performance depends primarily on capturing the persistence and nonlinear dynamics most relevant at each horizon.

Rehim Kılıç · 0 citations
Open access Sep 2026

Forecasting Inflation with Machine Learning and Traditional Time-Series Models: A Multi-Horizon Rolling-Origin Assessment

Introduction: The role of inflation forecasting in the monetary-policy assessment, financial planning and macroeconomic decision-making is crucial. This study also compares the Seasonal Autoregressive Integrated Moving Average (SARIMA) and Extreme Gradient Boosting (XGBoost) models to forecast Sticky Price Consumer Pri...

Shaista Sabir · 0 citations
Open access Aug 2026

Forecasting Multivariate Time Series: A Comparison of Machine Learning, Statistical and Deep Learning Models

The findings demonstrate that rigorous leakage-free validation is essential for reliable forecasting research and that, for monthly Robusta coffee prices, increased model complexity does not necessarily yield superior predictive performance.

Dler H Kadir, D. Khalil, Azhin M. Khudhur · 0 citations
Preprint Sep 2026

Does Training on Future Data Pay? Look-Ahead Bias in Forecasting with Pretrained Models

We examine whether post-origin training information inflates the measured accuracy and economic value of financial forecasts. We evaluate five sets of financial time-series foundation models, each comprising independently trained annual vintages under U.S., global, and factor-augmented training environments, across 14...

Hai-Qiang Chen, Li Chen, Yun-Long Chen et al. · 0 citations
Sep 2026

Models for real-time nowcasting and forecasting of GDP components by expenditure

Which family of models and which particular models are the most accurate for relatively short time series are asked, which approaches should be developed further, and why machine learning still yields no substantial gain over advanced econometric methods are compared are compared.

N. Fokin · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.