Adversarially Robust Self-Distillation
Abstract
To mitigate the vulnerability of deep neural networks in the face of adversarial attacks, a variety of defense strategies have been proposed in recent years. Adversarially robust distillation provides an effective approach by transferring knowledge from a robust teacher model to a student model. However, most existing robust distillation methods rely on a sufficiently robust teacher model, which is expensive to train in practice. This strong dependency greatly limits the practicality and scalability of robust distillation. To address this limitation, we propose an adversarially robust self-distillation (ARSD) method that does not require a robust teacher model as a training prerequisite. Starting from a standardly trained model, the proposed ARSD method guides adversarial training by leveraging the model’s own prediction. Besides, experimental results demonstrate that the performance of robust distillation is closely related to the divergence measure adopted in adversarial training. With reverse Jensen–Shannon divergence, the proposed ARSD method not only improves adversarial robustness, but also preserves a relatively high clean accuracy.