Skip to content
Open access

DWAT: Density-Weighted Adversarial Training for Robustness Beyond the Training Perturbation Budget

Sep 2026 · Applied Sciences · 0 citations

Abstract

Deep neural networks (DNNs) are widely deployed in safety-critical applications such as medical diagnosis and autonomous driving. Adversarial training (AT) is among the most effective defenses, casting robust optimization as a min–max problem over a defender-specified ℓp-ball of fixed radius ϵ. Bounded defenses of this kind are known to generalize poorly to test-time perturbations larger than ϵ. In this paper, we revisit this failure mode on MNIST and FashionMNIST and make three of its properties explicit. The degradation is abrupt rather than gradual, and its location is indexed by the training budget, so that enlarging the budget translates the drop instead of removing it. The translation is, in turn, capped by trainability, since training stops converging once ϵ becomes too large. Decision-surface visualization exhibits the same failure geometrically, as adversarially trained models form a plateau whose edge coincides with the training boundary. Together, these properties suggest that the failure follows from concentrating training on a single radius. We therefore present Density-Weighted Adversarial Training (DWAT), a plug-in framework that spreads training over a set of sampled ϵ-balls and reweights each candidate adversarial example by a Gaussian density of its distance from the benign sample. We derive its objective as a self-normalized importance-sampling estimate of an expected adversarial risk under a perturbation prior, and we show that this risk is upper-bounded by the standard adversarial risk, so that the in-bound objective of DWAT relaxes rather than replaces that of AT. We further prove that the population objective upper-bounds the worst-case adversarial risk at every radius, including radii beyond the training budget, at an explicit cost that grows with the radius. Experiments with PGD-AT, TRADES, and MART as base defenses indicate that DWAT alleviates the out-of-bound degradation in the settings that we study, at a cost inside the training budget that we report and discuss.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.