JevAdvBench is introduced, to the authors' knowledge the first adversarial benchmark for RLCD models, with 812 typed questions over 66 scenarios, and a black-box attack suite of 9,744 single-edit variants that each edit one part of a request, with billed input tokens confirming that the edit reached the model.
ODPure is proposed, a novel input-stage black-box defense for object detection, which is based on input purification that ensures stable perception flows and provides robust defense against diverse backdoor attacks and trigger types while preserving baseline accuracy.
Li Zeng, Ming-Cheng Duan, Long-Fei Fan et al.· 0 citations
This paper introduces the concept of instruction-dense visual jailbreaks, in which image-generation models produce detailed, readable, and actionable harmful instructions within images, and proposes TYPO, a black-box framework that exploits this safety gap by automatically generating adversarial TYPOgraphy prompts.
Meng Xie, Li Zeng, Hang Zhang et al.· arXiv.org· 0 citations
Vision-Language Models (VLMs) are known to be vulnerable to adversarial attacks, where subtle perturbations to images or texts induce erroneous outputs. However, most text-based attacks are adapted from language-model-centric methods, in which the visual input is fixed during optimization, resulting in adversarial prom...
Li Zeng, Ze-Yu Ye, Meng Xie et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.