Skip to content
Conference

A Three-Stage Residual-Boosted and Calibrated Framework for Multivariate Road Accident Risk Scoring

Aug 2026 · 2026 International Conference on Modern Sustainable Systems (CMSS) · pp. 1552-1558 · 0 citations · 20 references

Abstract

The conventional non-parametric machine learning models, such as Gradient Boosted Decision Trees (GBDTs), can be very powerful in terms of predictive capabilities for road traffic accident (RTA) risk indexing, but are vulnerable with respect to some architectural aspects: they lack deterministic domain-safety heuristics, they can produce out-of-bounds tail predictions, and they are cross validation unstable. In order to address these drawbacks, this paper presents a novel, Three-Stage Residual-Boosted and Calibrated Machine Learning Pipeline. Stage 1 is for macro level statistical distributions and regional roadway trends using a leaf-wise LightGBM regressor. Stage 2 incorporates an inductive engineering bias by defining a kinematic domain-safety prior and training an explicit XGBoost regressor that is only on the residual variance corrected by the prior. Stage 3 combines these complementary outputs in a Non-Negative Least Squares (NNLS) meta-learner with strict non-negativity and convex constraints, capped by a mathematical cap. Empirical validation shows that the pipeline performs an excellent linear alignment with real world ground truth with Pearson correlation coefficient out-of-fold $\mathbf{r}=\mathbf{0. 9 4 2}$. The framework is shown to effectively prioritize compound geometric-kinematic vectors (such as curvature-to-speed ratios) in addition to the engineered risk prior, through feature importance analysis. Moreover, NNLS optimization adds a dominant fractional weight of 0.62 to the Stage 2 residual learner, demonstrating that reducing the generalization error is more on the localized and highdimensional residual signals, but not on the traditional single-stage ensembling. The uniform error reduction and the stability of the weights for the meta-learner against data perturbations in stratified cross validation paths and in bootstrapping simulations are confirmed, and the boundary constraints are used to keep all risk indices in the physically valid range [0, 1] of the spectrum. Finally, this framework may help to join the gap between pure data-driven processes and traditional traffic safety heuristics, and create an extremely reliable and interpretable decision support tool for network-wide safety prioritization in ITS.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.