2026
Random Character-Level Perturbations Amplify LLM Jailbreak Attacks
This work finds that models cannot reliably reconstruct the original meaning and layer-wise probe classifiers fail to detect the harmful intent of perturbed prompts, and perturbations can occasionally reduce attack success by inducing off-topic or incoherent responses.
Shuyi Yu, Zhe Cao, Kohei Tsuji et al.
· Trans. Mach. Learn. Res. · 0 citations