Skip to content
Review

Vision-language and generative models in traffic video safety analysis: a computational framework and research agenda.

Jul 2026 · Accident Analysis and Prevention · Vol 236, pp. 108673 · 0 citations · 45 references
Medicine

Abstract

Vision-language and generative models have recently emerged as powerful tools for interpreting multimodal traffic-video data and advancing safety analysis. This paper reviews and integrates progress across foundation vision-language models, multimodal large language models, video-centric temporal reasoning frameworks, and diffusion-based world models, emphasizing how these paradigms enable richer semantic understanding, causal reasoning, and counterfactual safety assessment. We propose a unified taxonomy that maps model families to three application levels across diverse deployment environments-from cloud to onboard systems: network-scale monitoring, event-level crash and near-miss understanding, and generative or counterfactual scenario analysis. Key technical and methodological challenges are identified, including hallucination control, temporal consistency, sim-to-real transfer, and safety alignment with physical and rule-based constraints. The paper synthesizes open research problems and outlines a structured agenda toward grounded, interpretable, and computationally efficient multimodal cognition for real-world traffic safety applications.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.