AEGIS: A Mechanism-Guided Defense against Visual Synonym Jailbreaks in Text-to-Image Models
This work proposes AEGIS (Adaptive Evasion Guard via Identification and Steering), an inference-time defense that applies similarity-aware repulsion only at the identified vulnerable heads of VSA and preserves benign fidelity, avoids suppressing hard-negative concepts, and transfers to SD 2.1 and FLUX.