Vision-language and generative models in traffic video safety analysis: a computational framework and research agenda.
Vision-language and generative models have recently emerged as powerful tools for interpreting multimodal traffic-video data and advancing safety analysis. This paper reviews and integrates progress across foundation vision-language models, multimodal large language models, video-centric temporal reasoning frameworks,...