Public health forecasts must respond to abrupt changes in surveillance data without over-extrapolating noise, reporting artifacts, or temporary trends. We evaluated autoregressive integrated moving average (ARIMA), random forest, and extreme gradient boosting (XGBoost) models using 190 weekly observations of publicly available Ontario COVID-19 case counts from January 2020 to October 2023. Rolling-origin time-series cross-validation preserved temporal order during model tuning and evaluation. Performance was assessed across three operating dimensions: responsiveness following selected turning points, forecast horizons of one to six weeks, and the amount of historical training data. We also developed Machine Learning and ARIMA Model Averaging (MLAMA), a non-negative performance-weighted ensemble with weights that vary by forecast horizon and responsiveness setting. Retrospective comparisons showed that ARIMA adapted rapidly after turning points but its normalized error increased at longer horizons. Random forest and XGBoost were less responsive initially but maintained more stable normalized error over longer horizons. For two-week forecasts at the end of the study period, training on the most recent data outperformed using longer historical periods, particularly for XGBoost. MLAMA achieved the lowest normalized mean absolute percentage error across most forecast horizons and ranked among the best-performing methods across responsiveness settings. These findings support selecting forecasting models according to operating conditions rather than relying on a single universally preferred approach. MLAMA provides a practical framework for combining complementary statistical and machine-learning forecasts. The accompanying Python package is currently maintained in a private repository while software validation and reproducibility testing are completed.
In the energy transition of the world, the models to be used in power load prediction should be capable of delivering predictions that are not only accurate but also have a reasonable measure of uncertainty. The increase in the number of extreme weather events has caused the behavior of the loads to be nonlinear and unpredictable and this has restricted the effectiveness of the traditional deterministic forecasting approach in grid dispatching as well as warning of risk. To address pattern identification, data sparsity, and uncertainty under extreme weather, this paper develops an integrated probabilistic forecasting framework with three linked stages: extreme-weather load identification, TimeGAN-based sample augmentation, and conformal quantile forecasting. The method first builds a high-confidence extreme-weather load sample repository, then augments scarce extreme-weather sequences, and finally provides calibrated prediction intervals for short-term load forecasting. Experimental results show that the proposed method improves forecasting accuracy under extreme-weather conditions, with MAPE reduced across all five tested models after data augmentation; for example, ARIMA decreases from 12.37% to 8.65% and iTransformer decreases from 6.12% to 5.28%. The conformal quantile forecasting model also achieves 97.62% empirical coverage under the nominal 95% prediction interval, indicating improved prediction-interval reliability.
Hao Zhang, Xiyang Liu, Ruotian Gao et al.· International journal of pat...· 0 citations