Toward Sustainable Urban Mobility: A Multimodal Large Language Model (MLLM) Framework for Automated Driver Performance Assessment with YOLOv8-Based Scene Detection
An exploratory proof-of-concept framework for automated driver evaluation that combines real-world dashcam footage, YOLOv8-based object detection, and multimodal large language models (MLLMs), specifically Gemini 1.5 Flash is presented.
Abstract
Accurate and scalable driver performance assessment is critical for improving road safety and reducing traffic-related injuries and fatalities, particularly in low- and middle-income countries where the majority of global road deaths occur. This paper presents an exploratory proof-of-concept framework for automated driver evaluation that combines real-world dashcam footage, YOLOv8-based object detection, and multimodal large language models (MLLMs), specifically Gemini 1.5 Flash. Two prompting strategies, narrative and rule-based, were designed to assess driver behavior against standardized licensing criteria derived from the California Department of Motor Vehicles (DMV) driving performance evaluation score sheet. The framework was evaluated across 11 manually curated driving scenarios covering intersections, pedestrian crossings, stop signs, cyclists, and emergency vehicles. Ground-truth labels were established through consensus between two traffic engineering experts cross-referencing official California DMV evaluation criteria. In this preliminary evaluation, the rule-based prompt achieved higher agreement with ground-truth assessments (10/11 scenarios, 90.9%) compared to the narrative prompt (7/11 scenarios, 63.6%), particularly in detecting clear rule violations. The narrative approach demonstrated greater contextual flexibility in ambiguous situations. These results should be interpreted as preliminary, given the small sample size, manually curated dataset, and absence of large-scale statistical validation. Nonetheless, the findings illustrate how combining visual detection with structured language-model prompting may support interpretable, policy-aligned driver evaluation. Key limitations include dependence on video quality, limited scenario diversity, absence of temporal behavioral modeling, and reproducibility constraints tied to proprietary API behavior. Future work should expand validation to larger annotated datasets, incorporate temporal sequence modeling, and explore region-specific regulatory adaptation.
Abstract. This study develops an AI-driven platform to autonomously convert unstructured dashcam footage into structured data compatible with the Visual iRAP Data Analysis (VIDA) traffic management system. The platform utilizes YOLOv7 for object detection and vehicle trajectory extraction, Natural Language Processing (...
Safar Alhajri· Materials Research Proceedin...· 0 citations
Proper management of pedestrian flow at urban crosswalks is crucial for alleviating congestion, reducing waiting times, and improving general traffic safety. This paper presents AH-YOLOFlow, a computer vision framework for pedestrian detection, multi-object tracking, and zone-based flow analytics at urban crosswalks. T...
Farkhod Akhmedov, D. Khasanov, Sarvar Yusupov et al.· Sustainability· 1 citation
The safety and reliability of Automated Driving Systems (ADS) must be validated before large-scale deployment, and scenario-based testing is a promising way to improve validation efficiency and reduce cost. However, unidentified cross-regional differences in driving scenarios force manufacturers to repeat extensive...
Ji Zhou, Yongqi Zhao, Arno Eichberger· Journal of Intelligent and C...· 0 citations
Pedestrian detection and School-zone sign recognition are crucial aspect of advanced driver assistance systems and intelligent transportation systems. Consistent perception improves traffic safety and enhances the reliability of Advanced Driver Assistance Systems (ADAS). This research proposes a Spatial Attention-enhan...
M. Surya, Santanu Kumar Dash, N. Rajesh· IEEE Access· 0 citations
Assessing perceived road safety from street imagery remains a challenge because human judgments depend on infrastructure, context, and subjective interpretation. Current computer vision methods focus on object detection and semantic segmentation, with limited attention given to human-centric safety. This study prop...
Imad Tbaileh, Ahmed Radwan, Oroob Yaseen et al.· Frontiers in Future Transpor...· 0 citations
Driving assistance and autonomous driving are among the fastest-evolving domains of intelligent transportation systems (ITSs). However, visually impaired pedestrians (VIPs) remain weakly protected by roadsides or vehicle perception systems, especially when the right of way must be communicated in an accessible and audi...
Mourad Raif, A. El Rharras, Rachid Saadane et al.· Information· 0 citations
Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.