Skip to content

AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding

Jul 2026 · arXiv.org · Vol abs/2607.08745 · 0 citations · 12 references
Computer Science

TL;DR

The AUTOPILOT-VQA benchmark provides a standardized benchmark for assessing the reliability of autonomous driving systems in different scenarios and support developments for more interpretable, robust, and safety-conscious vision-language systems for real-world autonomous driving.

Abstract

Recent advances in Vision-Language Models, Large Language Models, and Multimodal Large Language Models have improved autonomous driving tasks such as scene understanding, decision making, trajectory prediction, and visual question answering. However, evaluating whether these models can reliably reason about safety-critical incidents remains challenging. To address this gap, we present AUTOPILOT-VQA, an incident-centric visual question answering benchmark for dashcam video understanding. The dataset evaluates different systems through structured questions designed around real-world driving incidents and near-incidents. The benchmark covers diverse safety-relevant categories, including weather and lighting conditions, traffic environment, road layout, road surface state, signage, involved entities, accident occurrence, impact location, and avoidability-related reasoning. By requiring models to answer grounded questions about both contextual scene properties and event-level incident details, AUTOPILOT-VQA moves beyond object recognition toward temporally grounded, safety-aware reasoning. The dataset is released as part of the AUTOPILOT CVPR 2026 competition and provides a standardized benchmark for assessing the reliability of autonomous driving systems in different scenarios. Our benchmark support developments for more interpretable, robust, and safety-conscious vision-language systems for real-world autonomous driving.

View source

Similar papers

Preprint Aug 2026

CAViAR: A Causal Video Dataset for Fine-Grained Accident Reasoning in Real-World Scenarios

CAViAR exposes a practical Perception--Reasoning Gap: current VLMs may recognize salient context, but do not reliably map visible agent actions to annotated rule-relevant responsibility categories in safety-critical driving scenarios, and all models degrade sharply on accident type and responsibility reasoning.

Sparsh Garg, Yi-Wen Chen, Vijay Kumar et al. · 0 citations
Jul 2026

ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness

This work introduces a real-world multi-modal benchmark for adverse-weather autonomous driving, designed with three capability dimensions: observability awareness, spatial reliability, and risk-aware decision-making, enabling fine-grained diagnosis of model behavior under degraded observations.

Qiao Yan, Yihan Wang, Zhenghao Xing et al. · 0 citations
Preprint Aug 2026

Inter-3D VQA: A Roadside Multimodal Benchmark for 3D Spatiotemporally Grounded Visual Question Answering

Recent advances in visual question answering (VQA) and multimodal large language models (MLLMs) have enabled natural-language reasoning over traffic scenes. However, existing benchmarks are largely built from ego-vehicle views or 2D roadside videos, limiting their ability to evaluate 3D-grounded reasoning over real-wor...

Shaozu Ding, Linan Song, Dajiang Suo · 0 citations
#artificial intelligence Preprint Sep 2026

SAFIRE: Safety-Critical Benchmark for Fine-grained Fire and Smoke Understanding in Multimodal LLMs

Multimodal Large Language Models (MLLMs) show strong progress on vision-language tasks, yet their reliability in safety-critical settings remains underexplored. Fire-smoke understanding is central to public safety and disaster response, but most existing benchmarks lack diverse real-world scenarios and context-aware ev...

Pengfei Li, Naufal Suryanto, Sicheng Zhang et al. · 0 citations
Preprint Aug 2026

RSVideo: Are Your Vision-Language Models Ready for Remote Sensing Videos?

Remote-sensing videos enable real-time observation of changes in target attributes, short-term activities, and scene evolution. They record motion, actions, interactions, and scene changes that cannot be captured by isolated images. Existing models primarily target single images or discrete temporal observations spanni...

Hongjie Zhou, Shiqin Wang, Haoyang Chen et al. · 0 citations
Preprint Aug 2026

UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations

UniTraffic-Agent is introduced, the MR-CAS solution for Track~3 of the 10th AI City Challenge, which includes Traffic Anomaly Reasoning (TAR) and two out-of-domain evaluations: FETV for fisheye traffic events and PSI-VQA for pedestrian intention reasoning.

Peng Li, Qianqian Xu, Shilong Bao et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.