Statistical and Economic Comparisons of Econometric, Linear, and Machine Learning Models
Abstract
This study compares econometric, linear regression, and machine-learning models for forecasting five-day close-to-close volatility across 12 diversified US exchange-traded funds. The dataset covers 2,511 trading days from September 2016 to August 2026 and includes equities, bonds, precious metals, commodities, and real estate. Fourteen models were evaluated using a leakage-free expanding-window framework, with quarterly retraining and hyperparameter tuning based only on training data. Performance was assessed using forecast errors, QLIKE, out-of-sample R-squared, statistical tests, the Model Confidence Set, regime analysis, SHAP explanations, and volatility-targeting outcomes. EGARCH achieved the best average QLIKE (0.382) and model rank (1.92), improving QLIKE by 12.2% over historical volatility and 5.3% over GARCH. The ensemble produced the lowest MAE (0.0412) and RMSE (0.0612) and the highest out-of-sample R-squared (0.159), while HAR achieved the highest volatility-targeting Sharpe ratio (0.895). Model rankings differed significantly (p < 0.001). Although EGARCH generally performed strongly, it did not significantly outperform GARCH or GJR-GARCH. Overall, machine-learning models did not consistently outperform econometric approaches, and the best model depended on the asset class, market conditions, evaluation measure, and forecasting purpose.