Temporal Cross-Modal Alignment Attack against Perception-Oriented Vision-Language Models for Autonomous Driving
Vision-language models (VLMs) have shown strong performance in autonomous driving (AD) tasks, supporting scene understanding and safety-related multimodal reasoning. However, robustness under adversarial perturbations remains critical, and the alignment vulnerability between visual evidence and task semantics under seq...