This paper introduces responsibility distribution estimation for ego-view traffic accident videos, a new task in which a model predicts the percentage of responsibility assigned to each involved agent, and constructs an LLM-assisted responsibility annotation pipeline and fine-tune multimodal large language models under multiple input settings.
Abstract
Recent studies on multimodal traffic accident understanding have mainly relied on infrastructure-camera footage, satellite imagery, or structured crash records. However, such data sources are costly to deploy and maintain at large scale, and they cannot objectively capture what the driver was actually able to observe before the accident. In contrast, ego-view accident videos directly represent the driver's visual perspective, making them suitable for reasoning about avoidability and driver responsibility. In this paper, we introduce responsibility distribution estimation for ego-view traffic accident videos, a new task in which a model predicts the percentage of responsibility assigned to each involved agent. We construct an LLM-assisted responsibility annotation pipeline and fine-tune multimodal large language models under multiple input settings, including raw frames, segmentation-enhanced input, and textual descriptions. Experimental results establish a strong initial benchmark, demonstrating that multimodal LLMs can effectively perform this nuanced, constraint-based reasoning task. Our findings suggest that ego-centric accident videos provide a promising foundation for socially and legally meaningful multimodal reasoning beyond conventional accident classification and explanation tasks.
CAViAR exposes a practical Perception--Reasoning Gap: current VLMs may recognize salient context, but do not reliably map visible agent actions to annotated rule-relevant responsibility categories in safety-critical driving scenarios, and all models degrade sharply on accident type and responsibility reasoning.
Sparsh Garg, Yi-Wen Chen, Vijay Kumar et al.· 0 citations
As road traffic accidents are high-frequency incidents. Although there are sufficient and clear traffic control cameras at some perception-dense sites, in other places there is often a lack of enough cameras. In order to make full use of limited information to efficiently obtain an overview of the accident scene in a s...
UniTraffic-Agent is introduced, the MR-CAS solution for Track~3 of the 10th AI City Challenge, which includes Traffic Anomaly Reasoning (TAR) and two out-of-domain evaluations: FETV for fisheye traffic events and PSI-VQA for pedestrian intention reasoning.
Peng Li, Qianqian Xu, Shilong Bao et al.· 1 citation
EgoSafe-Bench is introduced, a benchmark specifically designed to probe forensic reasoning in egocentric safety scenarios, generated by pairing each of the 3,000 video clips with a QA chain governed by the proposed Hierarchical Reasoning Evaluation (HRE) protocol.
A unified multi-attribute framework based on a Vision Transformer, which is enhanced with lightweight, parameter-efficient adapters and exhibits consistent behavior–context relationships and demonstrates robustness under varied environmental conditions is proposed.
Traffic surveillance cameras capture accidents continuously, yet converting raw CCTV footage into structured event records that pinpoint when, where, and what type of collision occurred remains unsolved at scale. The ACCIDENT @ CVPR benchmark evaluates exactly this joint prediction under a strict constraint: no labeled...
Dipit Saha, Shahruz Mannan, Mohammad Raihan Rashid et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.