Why Are LLM Backdoor Defenses Fragmented? A Feature-Level Explanation with Sparse Autoencoders
Results show that SAE-based analysis can explain defense fragmentation and guide interpretable backdoor mitigation and show system?atic encoding differences: dirty-label backdoors are dominated by isolated interaction features, whereas clean-label backdoors rely more on heterogeneous mixtures of mixed and weight-modifi...