Defending LLMs against Jailbreak Attacks via Template-Based ICL with a Defensive Suffix
State-of-the-art large language models (LLMs) have achieved impressive results on various tasks. However, these architectures are vulnerable to jailbreak attacks, such as GCG and Auto-DAN. Several defense strategies have been proposed to protect LLMs from generating harmful content, with most methods focusing on model...