Skip to content
#small language model Open access

Slow Drift Temporal Poisoning attacks and vision language model guided defense for BEV perception in autonomous vehicles

Sep 2026 · Discover Vehicles · Vol 2 · 0 citations · 32 references

TL;DR

Results indicate that temporal memory is an attack surface that BEV security evaluation has not yet accounted for, and that cross-modal semantic verification is a promising, if not yet field-validated, direction for defending it.

Abstract

World-model BEV perception maintains a persistent temporal latent memory that can be gradually corrupted by perturbations too small to trigger existing per-frame anomaly checks. We formalize this as Slow-Drift Temporal Poisoning (SDTP): frame-wise perturbations bounded by a small digital L∞ budget (ε ≤ 2/255) - a standard proxy for low perturbation magnitude, not a validated measure of human-perceptual imperceptibility or physical-world realizability, jointly optimized across a temporal window via a Backward Temporal Gradient procedure, that reach an 84.3% Temporal Drift Success Rate by frame 32 on BEVFormer-T while evading temporal-consistency defenses calibrated on single-step thresholds. We propose SemantiGuard, a cross-modal defense that cross-checks BEV detections against a frozen vision-language model’s scene description and triggers latent-memory rollback on divergence, reducing the Temporal Drift Success Rate (TDSR) from 91.2% to 13.8% under the primary evaluation configuration (SDTP-L: ε = 4/255, T = 32, BEVFormer-T, Table 6), with a lower 11.2% TDSR observed under the smaller SDTP-M ablation configuration (ε = 2/255, T = 16, Table 7). We also introduce TemporalAdvBEV, a benchmark extending RoboBEV with multi-frame adversarial configurations and drift-specific metrics. These results indicate that temporal memory is an attack surface that BEV security evaluation has not yet accounted for, and that cross-modal semantic verification is a promising, if not yet field-validated, direction for defending it. The paper also introduces TemporalAdvBEV, a benchmark extending RoboBEV with multi-frame adversarial configurations and drift-specific metrics, designed to support evaluation on both nuScenes and Waymo Open, with nuScenes results reported in this manuscript.

Read PDF

Similar papers

#computer vision Preprint Aug 2026

What's the Catch? Evaluating Temporal Consistency in Vision-Language Models

It is indicated that current VLMs can identify anomalies within individual frames but struggle to integrate information across frames to reason about temporal consistency, and TimeCatch provides a controlled benchmark for evaluating temporal grounding in vision-language models.

Marek Hradil, Danae Sánchez Villegas · 0 citations
Preprint Aug 2026

CertVLA: Certified Defense against Physical Visual Attacks for Vision-Language-Action Models

Vision-Language-Action (VLA) policies are vulnerable to localized physical perturbations, yet existing certified patch defenses target discrete labels and cannot directly certify continuous, temporally correlated actions. We introduce CertVLA, a certified defense for closed-loop VLA control under bounded patch and text...

Hui Lu, Zhi-Jie Peng, Yuqi Lin et al. · 0 citations
Preprint Aug 2026

Distilling Vision-Language Models for Robust Traffic Sign Perception in Autonomous Vehicles

Evaluated on GTSRB and LISA across four backbones and three physical attack types, LAMDA is the only method among ten evaluated that consistently improves robustness across all attack-backbone-dataset combinations, while preserving or improving clean accuracy in nearly all cases.

Pedram MohajerAnsari, Amir Salarpour, M. Pesé · 0 citations
Open access Aug 2026

Sparse Adversarial Patch Attack and Robustness Evaluation Algorithm for Vision-Language Models

Visual language models (VLMs) have demonstrated outstanding performance in high-value domains such as autonomous driving, unmanned system navigation, and intelligent question-answering; however, the security of their cross-modal alignment mechanisms has not yet been fully verified. Existing visual adversarial patch att...

T.-Y. Chen, X.-Y. Hu, J.-F. Wang et al. · 0 citations
Preprint Aug 2026

Glance, Scrutinize, and Think: Advancing Video Anomaly Detection from Training-Free to Agentic Reasoning

Glance then Scrutinize (GtS), a training-free framework using static and dynamic textual guidance for coarse-to-fine anomaly grounding and understanding, balancing accuracy and speed, and JeAUG, a metric jointly evaluating semantic interpretability and temporal precision are introduced.

Shibo Gao, Pei-Pei Yang, Xu-Yao Zhang et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Cog-VADU: A Training-Free Cognitive Reasoning Framework for Video Anomaly Detection and Understanding

Video Anomaly Detection (VAD) aims to temporally localize abnormal events in videos. Most existing approaches rely on dataset-specific training and curated annotations, limiting generalization in open-set scenarios. Recent zero-shot methods based on Large Vision- Language Models (LVLMs) alleviate this dependency but of...

Mohd Ubaid Wani, Sara Atito, Josef Kittler et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.