A context-aware and human-centered framework for road safety assessment using semantic scene understanding
Abstract
Assessing perceived road safety from street imagery remains a challenge because human judgments depend on infrastructure, context, and subjective interpretation. Current computer vision methods focus on object detection and semantic segmentation, with limited attention given to human-centric safety. This study proposes a Context-Aware Road Safety Assessment (CARSA) framework based on semantic scene understanding and multidisciplinary expert evaluation. A semantically guided image selection pipeline was applied to the Mapillary Vistas dataset to construct a representative set of road scenes. Following manual filtering, the images were classified into urban, residential, and rural contexts. A literature-informed, rule-based Context-Aware Road Safety Assessment (CARSA) framework was then defined using semantic scene features, predefined context-specific factor weights, and categorical decision thresholds. All framework parameters were fixed and CARSA predictions were generated before the expert annotation process. The images were subsequently evaluated independently by experts in transportation engineering, driving instruction, and computer vision, and the resulting expert consensus was used exclusively as an independent Ground Truth for paired human–framework agreement evaluation on the current benchmark. The expert evaluation demonstrated excellent reliability, achieving an ICC(A,k) of 0.953, a Cronbach’s Alpha of 0.954, and a Fleiss’ Kappa of 0.805. CARSA achieved an accuracy of 92.38%, a weighted F1 score of 92.61%, a Cohen’s Kappa of 0.853, and a Quadratic Weighted Kappa of 0.891. The non-contextual semantic baseline achieved 75.87% accuracy, whereas context-aware CLIP zero-shot, Qwen2.5-VL zero-shot, and Qwen2.5-VL few-shot achieved 19.68%, 20.32%, and 29.84%, respectively. CARSA significantly outperformed all four baselines in paired exact McNemar tests after Holm correction. Most CARSA disagreements were limited to adjacent safety categories. The proposed framework provides a reliable and explainable approach for assessing perceived road safety by integrating semantic scene understanding with contextual reasoning. Furthermore, the open benchmark, expert annotations, and evaluation protocol establish a reproducible foundation for future research in interpretable transportation systems, human–AI alignment analysis, and road scene understanding.