Skip to content
Preprint

Adversarial Training Without Input Gradients via Low-Rank Householder Expansions

Aug 2026 · 0 citations · 34 references
Computer Science

TL;DR

That such a regularizer exists is the main finding: the methods that dispense with the inner search all obtain their local geometry by differentiating with respect to the input, and it is shown this is not necessary.

Abstract

This work concerns adversarial training against the small-norm adversarial examples that arise from the inherent input instability of a trained deep neural network. Examples in this class are small as measured in the relative $\ell^2$-norm, and therefore lie in the neighborhood of the input on which the model acts approximately linearly, the regime in which the perturbation remains imperceptible. We first show that such examples can be computed directly from the trained network parameters, without input gradient iterations, by means of a linearization called the low-rank Householder expansion (LRHE). The expansion describes the composed affine map rather than any individual layer, and the directions it identifies are read from the activation pattern already available in the forward pass. We then propose a simple adversarial training scheme built on this construction. No differentiation with respect to the input is performed at any point: training requires only additional forward evaluations, with weight parameters updated by the standard backward pass, and the inner maximization of the usual min-max formulation is eliminated entirely. That such a regularizer exists is our main finding: the methods that dispense with the inner search all obtain their local geometry by differentiating with respect to the input, and we show this is not necessary. The regularizer costs the equivalent of $2.8$ PGD steps per epoch, an $8.7\times$ reduction relative to 40-step adversarial training on MNIST and below the cost of 3-step training. The resulting models match three-step PGD adversarial training for relative $\ell^2$ budgets $\varepsilon \le 0.02$ and 40-step training for $\varepsilon \le 0.012$, falling away beyond, consistent with the locality of the expansion.

View source

Similar papers

Open access Sep 2026

DWAT: Density-Weighted Adversarial Training for Robustness Beyond the Training Perturbation Budget

Deep neural networks (DNNs) are widely deployed in safety-critical applications such as medical diagnosis and autonomous driving. Adversarial training (AT) is among the most effective defenses, casting robust optimization as a min–max problem over a defender-specified ℓp-ball of fixed radius ϵ. Bounded defenses of this...

Jie-Ying Huang, Rui-Ming Zhu, Jia Xu et al. · 0 citations
Jul 2026

Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates

This work comprehensively investigates computation-efficient strategies to speed up latent adversarial training from two complementary perspectives, and reduces per-step adversarial-training FLOPs by 48.1% while requiring only 0.0118% trainable parameters.

Weiyi He, Yuping Lin, Jiliang Tang et al. · 0 citations
Open access Aug 2026

Analysis of adversarial examples in neural network image classifiers

This work explores different neural network architectures, including fully connected networks, classical convolutional networks, and residual networks, under four types of adversarial attacks constrained by different L p norms, and investigates how adversarial examples affect the internal representations of networks...

Jana Poľašková, Iveta Bečková, Stefan Pócos et al. · 0 citations
Conference Open access Sep 2026

PILO: Principal Component-based Implicit Regularization with Low-rank Optimization for Robust Transfer Learning

PILO is established, a new, more effective paradigm for robust transfer learning through principled and targeted parameter optimization, and significantly outperforms state-of-the-art full-parameter and parameter-efficient methods in robust accuracy across multiple benchmarks.

Shuaihe Liu, Qiugang Zhan, Guisong Liu et al. · 0 citations
Preprint Aug 2026

Continuous Adversarial MeanFlow Transfer

This work proposes MeanFlow-Transfer, which maps heterogeneous source outputs into a shared velocity representation, uses it to initialize an MF generator from the source weights, and optimizes an MF objective on the target domain, and introduces Continuous Adversarial MeanFlow, a post-training stage that extends conti...

Yara Bahram, Zahra Dehghani, M. Desbos et al. · 0 citations
Jul 2026

Adversarially Robust Self-Distillation

To mitigate the vulnerability of deep neural networks in the face of adversarial attacks, a variety of defense strategies have been proposed in recent years. Adversarially robust distillation provides an effective approach by transferring knowledge from a robust teacher model to a student model. However, most existing...

Rihao Li, Ran Wang, Meng Hu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.