LightEMMA: A Longitudinal Evaluation of Vision-Language Models for Autonomous Driving
Empirical findings show that, despite increased model scale and enhanced general reasoning capabilities, successive VLM generations do not consistently achieve better driving performance, and highlight the need for domain-specific adaptation to improve the safety of VLM-based autonomous driving systems.