Vision-language models (VLMs) are increasingly used in safety-related settings; therefore, it is vital to understand their underlying vulnerability to adversarial manipulation. Adversarial patches are practical local perturbations. However, targeted patch attacks on Contrastive Language-Image Pre-training (CLIP) remain insufficiently understood under strict spatial budgets and realistic imaging variation. Accordingly, we propose a target-oriented adversarial patch framework that addresses both patch placement and patch optimization. For patch placement, we introduce heatmap-guided localization including Random, Target Saliency, Gradient-weighted Class Activation Mapping (GradCAM), and Grad-Attention, and compare their performance. For patch optimization, we combine target cross-entropy, target-text feature alignment, competing-class suppression, original-class suppression, and total-variation regularization with an expectation-over-transformation (EOT) strategy. During optimization, random transformations such as rotation, perspective, color changes, and noise are applied to improve patch stability under varying imaging conditions. Experiments on OpenAI CLIP, OpenCLIP, MetaCLIP, and EVA-CLIP show that the three semantic guidance strategies improve attack success over random placement, and among them, Grad-Attention performs best on average because it aligns more closely with the model’s internal attention structure. Ablation studies on patch size, loss composition, and target class further show that cross-entropy is the most critical loss term and semantically distinctive targets are easier to induce. Furthermore, cross-model evaluation indicates that patches exhibit non-trivial transferability within the CLIP family, particularly when EVA-CLIP is used as the source model.
Xue-Ying Wang, Ding-Yi Lu, Cheng-Ci Hu et al.· IEEE Access· 0 citations
Infrared vision-language models (IR-VLMs) have emerged as a promising paradigm for multimodal perception under low-visibility conditions, yet their robustness to targeted adversarial attacks remains poorly understood. Existing adversarial patch methods mainly study RGB-based models or a single downstream task and do not characterize whether localized perturbations can induce an intended semantic target in IR-VLMs. We propose InfraPatch, a white-box, per-instance framework for targeted digital grayscale patch attacks against IR-VLMs. InfraPatch optimizes a compact single-channel patch within an approximately 5% local-area budget, combines proxy-guided placement with task-adaptive semantic objectives, and induces target behaviors in image classification, image captioning, and binary visual question answering. We evaluate ten infrared-adapted model variants on 300 synthetic infrared-style images generated by applying DiffV2IR to a fixed 30-category COCO subset, using clean-conditioned targeted success criteria. InfraPatch achieves targeted attack success rates from 86.00% to 100% across the ten variants. On CLIP and BLIP-2, proxy location search improves success by 6.67 and 10.33 percentage points over optimized random placement, respectively; LLaVA-1.5 remains saturated near 100% under both settings. Patch-area and objective ablations further expose substantial differences in vulnerability across architectures and task formats. These results show that small grayscale patches can inject chosen target semantics across IR-VLM families under a controlled digital threat model, motivating stronger robustness evaluation for infrared multimodal systems.
Chengyin Hu, Ding-Yi Lu, Jiajun Han et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.