Machine Learning and Spatial Context for Urban Road Traffic Crash Severity Prediction
Abstract
Objectives: This study examines how demographic, temporal, vehicular, and spatial conditions jointly influence severe road traffic crashes in a dense metropolitan road network and evaluates whether machine learning can support policy-relevant risk screening. Methods/Analysis: Police-based crash records from 2020-2024 were integrated with point-of-interest (POI), district, temporal, and vehicle variables. Crash severity was defined as a binary outcome, where severe/fatal crashes included at least one serious injury or death. Logistic Regression, Random Forest, and XGBoost were trained using stratified sampling, class weighting, and threshold optimization, and were assessed through ROC-AUC, PR-AUC, sensitivity, specificity, F1-score, and calibration. Findings: At the modeling-record level, severe/fatal crashes represented 78,038 cases (28.4%), whereas non-severe crashes represented 196,597 cases (71.6%). Random Forest achieved the strongest discrimination (ROC-AUC = 0.756; PR-AUC = 0.733) and the highest severe-case detection, while Logistic Regression provided the most transparent and better-calibrated probability estimates. Important predictors included age, time of day, motorcycle involvement, and proximity to gas stations, schools, convenience stores, restaurants, and other POIs. Novelty/Improvement: The study conceptualizes POIs as active urban exposure systems rather than passive location markers, while acknowledging that the observed associations do not establish causality. The findings support a dual-model strategy combining operational screening with interpretable policy analysis.