Aug 2026· 2026 IEEE International Conference on Cybernetics and Intelligent Systems (CIS) and IEEE International Conference on Robotics, Automation and Mechatronics (RAM)· pp. 307-312· 0 citations· 11 references
Abstract
Vision-language models (VLMs) have shown strong performance in autonomous driving (AD) tasks, supporting scene understanding and safety-related multimodal reasoning. However, robustness under adversarial perturbations remains critical, and the alignment vulnerability between visual evidence and task semantics under sequential observations is insufficiently explored. This paper proposes TCMA, a Temporal Cross-Modal Alignment Attack combining a task-oriented objective, an alignment disruption loss, and a lightweight temporal propagation mechanism to attack perception-oriented VLMs in AD. Specifically, TCMA constructs a semantic anchor from the task prompt to suppress correct visual-text alignment, and warm-starts each frame’s attack from the previous perturbation while enforcing temporal consistency. On BDD100K with Dolphins, TCMA achieves 50.0\% overall center-frame targeted ASR across traffic-light, pedestrian, and rider tasks. Transfer evaluation on Qwen2.5-VL further reaches 73.3\% center-frame and 76.7\% vote-level overall targeted ASR, demonstrating strong cross-model generalizability.
To efficiently generate adversarial examples, a gradient-based attack method is proposed that performs optimization exclusively on the vision encoder of the VLM rather than on the entire multimodal architecture, which significantly reduces the computational cost and resource requirements of the attack while maintaining...
Ilan Zini, B. Addad, Katarzyna Kapusta· 1 citation
Evaluated on GTSRB and LISA across four backbones and three physical attack types, LAMDA is the only method among ten evaluated that consistently improves robustness across all attack-backbone-dataset combinations, while preserving or improving clean accuracy in nearly all cases.
Pedram MohajerAnsari, Amir Salarpour, M. Pesé· 0 citations
Visual language models (VLMs) have demonstrated outstanding performance in high-value domains such as autonomous driving, unmanned system navigation, and intelligent question-answering; however, the security of their cross-modal alignment mechanisms has not yet been fully verified. Existing visual adversarial patch att...
T.-Y. Chen, X.-Y. Hu, J.-F. Wang et al.· Advanced Electromagnetics· 0 citations
Alignment-Guided Flow Transformer (AGFT) is presented, a novel framework that explicitly enforces tri-modal alignment through a dedicated alignment loss, bridging the representational gap across modalities and enhancing task adaptation.
Sheng-Chao Hu, Peng Wang, Qi-Yang Zhou et al.· 0 citations
A detection-guided dynamic attention steering system that leverages the locality insight from CNNs to efficiently steer a VLM’s attention toward more relevant sections of an image and demonstrates the effectiveness of combining different model architectures to harness their respective strengths for advancing VLM capabi...
Alan William Zhang, Rui Pan, Mike Wong et al.· 0 citations
Results indicate that temporal memory is an attack surface that BEV security evaluation has not yet accounted for, and that cross-modal semantic verification is a promising, if not yet field-validated, direction for defending it.
S. Hallur, G. Baskaran, Hari Krishna Kattoju et al.· Discover Vehicles· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.