Efficient Frame Retrieval for Traffic Law Question Answering over Dashcam Videos
Abstract
Traffic Law Question Answering over dashcam videos requires both accurate visual evidence selection and reliable multimodal reasoning. This task is especially challenging in Vietnamese traffic scenes, where road environments are dense, diverse, and highly dynamic. In this paper, we present a resource-efficient two-stage framework for traffic-law question answering over dashcam videos, developed for the "RoadBuddy: Understanding the Road through Dashcam AI" track of ZaloAI Challenge 2025. First, a CLIP ViT–based frame retriever reduces visual redundancy by selecting four query-relevant frames from each video. Second, the selected frames are processed by a Qwen3-VL 8B model adapted with LoRA for task-specific multiple-choice answer prediction. Despite its lightweight design, our method achieves competitive results, with scores of 0.70617 on the public leaderboard and 0.702 on the private leaderboard, ranking top 3 among 241 participating teams. These results show that efficient frame retrieval, when combined with parameter-efficient task adaptation, offers a practical solution for traffic-law question answering over real-world dashcam videos.