Unified Accident Anticipation and Explainable Report Generation via Self-Supervised Vision-Conditioned Language Modeling
Accident anticipation from dashcam videos is essential for advanced driver-assistance systems, yet existing methods often provide limited interpretability beyond a risk score. We propose a unified multi-task framework that jointly performs frame-level incident anticipation, accident category recognition, and structured...