How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE
Directional ablation removes an aligned language model's ability to refuse by projecting a single"refusal direction"out of the weights that write the residual stream. It needs no gradient-based training and no optimization, only a few hundred contrastive prompts, which makes it the canonical white-box attack on open-we...