Skip to content
Conference

Temporal Cross-Modal Alignment Attack against Perception-Oriented Vision-Language Models for Autonomous Driving

Aug 2026 · 2026 IEEE International Conference on Cybernetics and Intelligent Systems (CIS) and IEEE International Conference on Robotics, Automation and Mechatronics (RAM) · pp. 307-312 · 0 citations · 11 references

Abstract

Vision-language models (VLMs) have shown strong performance in autonomous driving (AD) tasks, supporting scene understanding and safety-related multimodal reasoning. However, robustness under adversarial perturbations remains critical, and the alignment vulnerability between visual evidence and task semantics under sequential observations is insufficiently explored. This paper proposes TCMA, a Temporal Cross-Modal Alignment Attack combining a task-oriented objective, an alignment disruption loss, and a lightweight temporal propagation mechanism to attack perception-oriented VLMs in AD. Specifically, TCMA constructs a semantic anchor from the task prompt to suppress correct visual-text alignment, and warm-starts each frame’s attack from the previous perturbation while enforcing temporal consistency. On BDD100K with Dolphins, TCMA achieves 50.0\% overall center-frame targeted ASR across traffic-light, pedestrian, and rider tasks. Transfer evaluation on Qwen2.5-VL further reaches 73.3\% center-frame and 76.7\% vote-level overall targeted ASR, demonstrating strong cross-model generalizability.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Breaking the weakest link to evade vision language models

To efficiently generate adversarial examples, a gradient-based attack method is proposed that performs optimization exclusively on the vision encoder of the VLM rather than on the entire multimodal architecture, which significantly reduces the computational cost and resource requirements of the attack while maintaining...

Ilan Zini, B. Addad, Katarzyna Kapusta · 1 citation
Preprint Aug 2026

Distilling Vision-Language Models for Robust Traffic Sign Perception in Autonomous Vehicles

Evaluated on GTSRB and LISA across four backbones and three physical attack types, LAMDA is the only method among ten evaluated that consistently improves robustness across all attack-backbone-dataset combinations, while preserving or improving clean accuracy in nearly all cases.

Pedram MohajerAnsari, Amir Salarpour, M. Pesé · 0 citations
Open access Aug 2026

Sparse Adversarial Patch Attack and Robustness Evaluation Algorithm for Vision-Language Models

Visual language models (VLMs) have demonstrated outstanding performance in high-value domains such as autonomous driving, unmanned system navigation, and intelligent question-answering; however, the security of their cross-modal alignment mechanisms has not yet been fully verified. Existing visual adversarial patch att...

T.-Y. Chen, X.-Y. Hu, J.-F. Wang et al. · 0 citations

Detection-Guided Attention Steering for Vision Language Models

A detection-guided dynamic attention steering system that leverages the locality insight from CNNs to efficiently steer a VLM’s attention toward more relevant sections of an image and demonstrates the effectiveness of combining different model architectures to harness their respective strengths for advancing VLM capabi...

Alan William Zhang, Rui Pan, Mike Wong et al. · 0 citations
#small language model Open access Sep 2026

Slow Drift Temporal Poisoning attacks and vision language model guided defense for BEV perception in autonomous vehicles

Results indicate that temporal memory is an attack surface that BEV security evaluation has not yet accounted for, and that cross-modal semantic verification is a promising, if not yet field-validated, direction for defending it.

S. Hallur, G. Baskaran, Hari Krishna Kattoju et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.