Aug 2026· Frontiers in Applied Mathematics and Statistics· Vol 12· 0 citations· 52 references
TL;DR
Hierarchical Bayesian smooth transition autoregressive models provide accurate forecasts by accommodating nonlinear regime-switching dynamics while delivering robust uncertainty quantification, which makes them well suited for infectious disease surveillance and public health decision-making in resource-limited, high-uncertainty settings.
Abstract
Public health policy and disease surveillance systems require accurate forecasting of infectious disease dynamics to support timely interventions and resource allocation. However, classical linear time-series models often fail to capture abrupt regime shifts, nonlinear transmission patterns, and heterogeneous reporting commonly observed in surveillance data.
This study investigates the practical advantages of hierarchical Bayesian smooth transition autoregressive (BH-STAR) models, including logistic (LSTAR) and exponential (ESTAR) specifications. Performance is evaluated through extensive simulation studies under controlled nonlinear data-generating mechanisms, followed by an empirical application to COVID-19 surveillance data from 53 African countries collected between March 2020 and December 2022.
Simulation studies revealed a key directional asymmetry in model misspecification: fitting a logistic transition function to data generated under exponential dynamics resulted in moderate, stable parameter bias, whereas fitting an exponential transition function to logistic dynamics induced severe, compounding bias. Despite this parameter confounding under model misspecification, predictive accuracy remained stable across both data-generating processes. In the empirical application, the BH-STAR models consistently outperformed linear alternatives in out-of-sample forecasting. The hierarchical logistic STAR (HLSTAR) model achieved the highest overall predictive accuracy, reducing validation errors to an MAE of 0.58 and a MdAPE of 12.8%, corresponding to country-level forecast error reductions of 30–50% and overall error reductions exceeding 50% relative to the standard autoregressive benchmark.
Hierarchical Bayesian smooth transition autoregressive models provide accurate forecasts by accommodating nonlinear regime-switching dynamics while delivering robust uncertainty quantification. These features make them well suited for infectious disease surveillance and public health decision-making in resource-limited, high-uncertainty settings.
Epidemic modelling increasingly combines routine surveillance, mechanistic transmission theory and data-intensive learning, yet the resulting approaches are often compared as though they estimate the same quantities and serve the same decisions. This critical narrative review evaluates statistical, mathematical, machine-learning and hybrid approaches to infectious disease dynamics, with emphasis on inferential purpose, data-generating and observation processes, uncertainty, validation, interpretability and public-health use. Peer-reviewed literature published from January 2000 to 31 May 2026 was identified through biomedical, multidisciplinary and computing-oriented scholarly sources, supplemented by citation searching and verification against authoritative bibliographic records. Foundational earlier papers were retained when necessary. The synthesis indicates that no model class is uniformly superior. Statistical surveillance and time-series models are often efficient for anomaly detection, nowcasting and short-horizon forecasting, but their parameters rarely support intervention counterfactuals without additional causal structure. Mechanistic compartmental, network, spatial and agent-based models make transmission assumptions explicit and can represent intervention pathways, although structural misspecification, weak identifiability and mismatch between latent infections and observed reports can dominate their uncertainty. Machine-learning models can extract nonlinear and high-dimensional patterns from heterogeneous data, but apparent accuracy is vulnerable to temporal or spatial leakage, changing surveillance systems, distribution shift, weak probabilistic calibration and limited causal meaning. Hybrid models can combine epidemiological constraints with flexible learning and data assimilation; their interpretability nevertheless depends on identifiable parameters, biologically coherent architecture and validation beyond the setting used for training. Across paradigms, the observation process, target definition and decision horizon are as consequential as model form. Credible use therefore requires question-first model selection, explicit separation of forecasts from scenarios, rolling and geographically external validation, calibrated uncertainty, versioned data and code, and transparent communication of assumptions. Progress will depend less on greater complexity alone than on prospective benchmarking, identifiable hybridisation, behaviour-aware causal designs, multimodal surveillance, equitable data systems and operational model governance.
Hamid H. Hussien, Muhammed Aljifri, Nuha Hassan Hagabdulla et al.· Asian Journal of Probability...· 0 citations
Time series forecasting is a pivotal tool across multiple disciplines, particularly in epidemiology, where precise predictions can guide resource allocation and inform public health strategies. This study explores the application of Seasonal Autoregressive Integrated Moving Average (SARIMA) models for short-term forecasting of time series data, using epidemiological data from Kazakhstan related to COVID-19 as a practical case study. The dataset encompasses total confirmed cases, daily ambulatory care patients, and daily hospitalized patients, spanning from May 30, 2022, to December 14, 2022, with forecasts extending 10 days forward to December 24, 2022. The analysis leverages SARIMA’s ability to capture both seasonal and non-stationary patterns, making it highly effective for modeling complex epidemiological dynamics.
The methodology involves several key steps: data preprocessing through natural-log transformation to stabilize variance, stationarity testing using autocorrelation function (ACF) plots, and differencing to address non-stationarity. The optimal SARIMA models identified were (5,3,1)(1,1,0)_7 for total cases, (1,2,2)(0,1,1)_7 for ambulatory care patients, and (2,2,1)(2,1,1)_7 for hospitalized patients, with forecasting accuracies of 99%, 97%, and 91%, respectively. These high accuracies, measured via Mean Absolute Percentage Error (MAPE), underscore SARIMA’s robustness in short-term predictions. Residual analysis, including Shapiro-Wilk and Ljung-Box tests, confirmed that the models’ residuals were normally distributed and independent, validating their suitability for forecasting.
This study highlights SARIMA’s versatility in capturing weekly seasonal patterns, as observed in the 7-day cycles within the COVID-19 data, which reflect reporting or behavioral trends. While the example focuses on epidemiological data, the methodology is broadly applicable to other domains, such as ecological monitoring or economic forecasting, where seasonal time series are prevalent. The findings demonstrate that SARIMA models provide a reliable framework for short-term forecasting, offering actionable insights for public health interventions and resource planning, with the COVID-19 dataset serving as an illustrative example of the approach’s efficacy and adaptability.
M. Sorokina, I. Korshukov, N. Omarbekova et al.· Medicine and ecology· 0 citations
Deterministic dynamical models are widely used in infectious disease modelling, but they often become systematically biased when key drivers such as seasonality and random environmental variation are simplified or omitted. This paper proposes a practical way to account for structural bias while preserving the underlying mechanistic model. We model the observed time series as the sum of (i) a deterministic transmission component given by a reduced Ross malaria model and (ii) a latent stochastic seasonal component that captures unresolved seasonal forcing and other unmodelled variability. The seasonal component is defined as a mean-reverting stochastic differential equation with periodic forcing, which can be interpreted as a seasonally forced Ornstein–Uhlenbeck (OU) process and, at the same time, as a dynamically constrained model-discrepancy term. The combined model forms an additive Bayesian state-space system. A key challenge in additive decompositions is identifiability: many combinations of the deterministic and stochastic components can explain the same observations. We therefore use informative priors to stabilise this decomposition, together with a non-centred parameterisation (a reparameterisation that improves MCMC efficiency by sampling standardised noise terms instead of states directly) of the latent stochastic differential equation (SDE) states that enables efficient joint inference with Hamiltonian Monte Carlo (NUTS) in Stan. We validate the approach in three steps: (1) an OU example that contrasts parameter-only inference with latent state-space inference, (2) synthetic experiments that have a good agreement with the observed signal while highlighting the expected negative posterior dependence between components, and (3) an application to five years of monthly malaria case reports from Delta State, Nigeria. In the real-data analysis, the proposed additive ODE–SDE model produces a close fit with coherent uncertainty quantification and a flexible seasonal reconstruction, while keeping the mechanistic transmission model interpretable. Overall, the framework provides a flexible and transferable method for accounting for structural model discrepancy in misspecified dynamical models using a structured stochastic discrepancy.
Miracle Amadi, J. García-Merino, H. Haario· Bulletin of Mathematical Bio...· 0 citations
Non-stationary time series are common in many real-world domains, including infectious disease spread, where the underlying relationships between variables evolve over time. However, most existing forecasting methods assume stationarity and fail to capture changing causal dynamics. To address this challenge, we propose the Causal Regime Bayesian (CaReBayes) forecasting framework, which integrates regime detection, causal discovery, and Bayesian forecasting within a unified approach. CaReBayes segments time series into regimes using temporal causal discovery, fits a Bayesian structural autoregressive model for each regime, classifies the current regime, then performs regime-specific forecasting with uncertainty quantification. The framework introduces methodological advances: a grid-search procedure for automated regime-dependent causal discovery, a classification method that assigns future observations to regimes based on learned Bayesian structures, and regime-conditioned Bayesian structural forecasting. Across both synthetic and Ontario COVID-19 time series data, CaReBayes outperforms benchmark models for time series forecasting. In addition to improved forecasting performance, it produces regime-dependent causal graphs that summarize candidate structural relationships in the system, enhancing interpretability.
Autocorrelated bivariate count data frequently arise in criminal, environmental and financial studies, where capturing both serial dependence and cross‐series interaction is essential for statistical modelling and inference. In many applications, such dynamics are further influenced by exogenous covariates such as policy interventions or environmental factors, leading to time‐varying dependence structures that are not adequately captured by standard models. Existing bivariate integer‐valued autoregressive (BINAR) models mainly rely on constant or observation‐driven coefficients and rarely incorporate covariate information, which restricts their ability to represent evolving dependence in multivariate count processes. To address this limitation, we propose a covariate‐driven doubly stochastic bivariate integer‐valued autoregressive process, in which the thinning mechanism evolves jointly with past observations and exogenous covariates. This formulation extends the classical BINAR framework by allowing the dependence structure to vary dynamically under both internal and external driving mechanisms. The basic statistical properties of the proposed process are derived, and two estimation methods are developed, including an EM‐based algorithm. Monte Carlo simulations and a real data application are conducted to assess finite‐sample performance and robustness under different settings.
Public health forecasts must respond to abrupt changes in surveillance data without over-extrapolating noise, reporting artifacts, or temporary trends. We evaluated autoregressive integrated moving average (ARIMA), random forest, and extreme gradient boosting (XGBoost) models using 190 weekly observations of publicly available Ontario COVID-19 case counts from January 2020 to October 2023. Rolling-origin time-series cross-validation preserved temporal order during model tuning and evaluation. Performance was assessed across three operating dimensions: responsiveness following selected turning points, forecast horizons of one to six weeks, and the amount of historical training data. We also developed Machine Learning and ARIMA Model Averaging (MLAMA), a non-negative performance-weighted ensemble with weights that vary by forecast horizon and responsiveness setting. Retrospective comparisons showed that ARIMA adapted rapidly after turning points but its normalized error increased at longer horizons. Random forest and XGBoost were less responsive initially but maintained more stable normalized error over longer horizons. For two-week forecasts at the end of the study period, training on the most recent data outperformed using longer historical periods, particularly for XGBoost. MLAMA achieved the lowest normalized mean absolute percentage error across most forecast horizons and ranked among the best-performing methods across responsiveness settings. These findings support selecting forecasting models according to operating conditions rather than relying on a single universally preferred approach. MLAMA provides a practical framework for combining complementary statistical and machine-learning forecasts. The accompanying Python package is currently maintained in a private repository while software validation and reproducibility testing are completed.