Controlled Decoding Attacks on Black-Box LLMs
This work introduces \method{}, a framework for jailbreaking through text-only continuation interfaces that permit repeated sampling and assistant-prefix continuation, and achieves the highest mean score most comparisons against baselines.
Jesson Wang, Shawn Li, Wei Yang et al.
· 0 citations