Skip to content

A noise-based defense for stealthy backdoor attacks in large vision-language models

TL;DR

Control noise injection is investigated as a lightweight input-side defense against BadVision-style backdoors and suggests that stealthy encoder-level triggers depend on fragilestatistical patterns and can be weakened through controlled noise injection without requiring training of the full multimodal model.

Abstract

Large vision-language models rely on pretrained vision encoders to translate images intofeature representations used by downstream language models. This creates a security riskwhen the encoder is compromised by a stealthy backdoor attack, such as BadVision, where asubtle trigger causes an image to be mapped toward an attacker-chosen target representationwhile clean inputs remain largely unaffected. Because the model behaves normally understandard evaluation, these attacks are difficult to detect.This thesis investigates controlled noise injection as a lightweight input-side defenseagainst BadVision-style backdoors. The proposed approach adds small perturbations toinput images before they enter the vision encoder, with the goal of disrupting the triggerwhile preserving the semantic content of clean images. Several perturbation types are evaluated, including Gaussian noise, random noise, salt-and-pepper noise, low-frequency noise,geometric transformations, occlusion, scaling, rotation, and channel-based distributions.Experimental results show that geometric and channel-based transformations have limitedeffect on the backdoor, while pixel-level statistical perturbations significantly reduce targetsimilarity, increase feature-space distance from the attacker’s target representation, and lowerattack success. These findings suggest that stealthy encoder-level triggers depend on fragilestatistical patterns and can be weakened through controlled noise injection without requiringretraining of the full multimodal model.

View source

Similar papers

Preprint Aug 2026

Adversarial Attacks on Deep OCR Systems

Deep-OCR (DeepSeek-OCR) advances document recognition by treating the visual modality as an optical compression medium, enabling long-context OCR at low token cost. However, its increased complexity may introduce new security vulnerabilities. In this paper, we present, to the best of our knowledge, the first pure black...

Wenbo Sun, Hong-Zong Li, Yanyun Wang et al. · 0 citations
Preprint Sep 2026

FreqDoor: A Hidden Trojan in the Frequency Domain for Backdoor Attacks on Vision-Language Models

Vision-language models (VLMs) have recently shown excellent progress in open-ended image-to-text generation. However, their multimodal nature makes them persistently vulnerable to backdoor attacks. Existing backdoor triggers for VLMs are either spatial, textual, or bimodal, which may yield localized or recognizable tri...

Unknown authors · 0 citations
Jul 2026

GhostPrompt: Cross-Image Adversarial Prompt for Vision-Language Models

Vision-Language Models (VLMs) are known to be vulnerable to adversarial attacks, where subtle perturbations to images or texts induce erroneous outputs. However, most text-based attacks are adapted from language-model-centric methods, in which the visual input is fixed during optimization, resulting in adversarial prom...

Li Zeng, Ze-Yu Ye, Meng Xie et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Discovering Natural Transformation Vulnerabilities in Black-Box Vision Models

Natural adversarial examples (NAEs) reveal that vision models can fail under realistic semantic changes beyond norm-bounded perturbations. However, generating NAEs in a black-box setting remains challenging because existing generative attacks often rely on surrogate models, learned attack priors, or costly query-based...

Dong-Su Song, Dae-Yoo Go, J. Jung · 0 citations
Preprint Aug 2026

DEFUSE: Generalizable Backdoor Defense for Self-Supervised Encoders with Generative Priors

Inspired by Bayesian posterior inference, this work reformulate backdoor detection as a representation-conditioned image likelihood estimation problem parameterized by a conditional diffusion generative model, and fine-tune a pretrained diffusion model, leveraging its generative prior to map data onto the natural image...

Tuo Chen, Jie Gui, Minjing Dong et al. · 0 citations
Conference Open access Sep 2026

Understanding and Exploiting Phase Sensitivity for Attacking Large Vision–Language Models

This paper proposes a novel LVLM attack method, called BadPhase with further backdoor designs, to implant adversarial phase as triggers into any image inputs via data poisoning so as to control the LVLMs’ predictions and finds that LVLMs are sensitive to the phase-aware image structure.

Dai-Zong Liu, Junhao Dong, Xiang Fang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.