Skip to content

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks

Jul 2026 · arXiv.org · Vol abs/2607.23015 · 0 citations · 30 references
Computer Science

TL;DR

Mask2Shield (M2S), a masked-forward alignment method that trains a model under this functional pruning procedure, reduces successful recomputed pruning attacks from 80--279 to 1--44 out of 313 prompts while generally preserving four capability benchmarks.

Abstract

Large language models (LLMs) are safety-aligned before deployment to reduce harmful content generation. Yet neuron-level pruning attacks show that refusal can depend on a small set of removable units: disabling them can remove safety behavior while leaving much of the model usable. To address this problem, we introduce Mask2Shield (M2S), a masked-forward alignment method that trains a model under this functional pruning. The masked student must recover a safe refusal through the remaining computation, while a frozen, unmasked teacher supplies complete benign answers to limit capability drift. Across ten model configurations, M2S reduces successful recomputed pruning attacks from 80--279 to 1--44 out of 313 prompts while generally preserving four capability benchmarks. We also evaluate M2S with TwinBreak, which uses a different neuron-selection rule and iterative pruning procedure. Together, these results show that M2S makes targeted pruning less effective by reducing reliance on a small, removable safety-neuron set.

View source

Similar papers

Jul 2026

Defense Against LLM Backdoors using Critical Neuron Isolation Pruning

DeCNIP (Defense with Critical Neuron Isolation Pruning), which leverages representational analysis to identify and neutralize backdoors in a unified pipeline, is introduced, which achieves over 95% relative reduction in Attack Success Rate (ASR), outperforming seven state-of-the-art defenses with only 0.1% neuron inter...

Yuxi Li, Zhi-Bo Zhang, Kailong Wang et al. · 0 citations
Preprint Aug 2026

Tripwire: Triggering Aligned Refusal via Statistically Certified Safety Neurons

Neuron- and path-level interventions offer the finest-grained route to defending large language models (LLMs) against jailbreak attacks, yet existing methods fall short of this promise, i.e., they often compromise model utility significantly. Specifically, one line of work suppresses toxic neurons to erase harmful sema...

Wei Zhao, Zhe Li, Pei-Xin Zhang et al. · 0 citations
Preprint Aug 2026

NeuronGuard: Robust LLM Safety Alignment via Ablation-Aware Safety Signal Redistribution

A fine-tuning-stage defense that simultaneously hardens LLMs against both attack classes by redistributing safety signals across a broader set of neurons, and provides a formal guarantee that NeuronGuard strictly reduces the attack success rate (ASR) upper bound.

Anjun Gao, Yueyang Quan, Yu Xia et al. · 2 citations
Preprint Sep 2026

AlcaTRAz - Anchored Tree-Rule Defense Against Jailbreaks

Large language models (LLMs) are vulnerable to jailbreak attacks that bypass safety alignment through carefully crafted prompts. Many existing defenses require access to model weights or internals, making them difficult to apply to black-box deployments. We propose AlcaTRAz (Anchored Tree-Rule defense Against jailbreak...

J. Res, Petr Kaska, Martin Perešíni et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Bait-and-Recover: Poisoning Internal Refusal Signals to Defend LLMs against White-Box Editing Jailbreaks

Open-weight large language models face a low-cost white-box threat from representation engineering attacks. Attackers can estimate refusal directions and search for projection-matrix edits that suppress safety alignment while preserving general capabilities, within minutes on a single GPU and without gradient-based tra...

Tian Gao, Zhi-Hui Xie, Yu-Hao Wu et al. · 0 citations
Preprint Aug 2026

NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation

NeuronFuzz is presented, a white-box fuzzing framework that exploits internal safety neurons as continuous execution feedback for LLM safety evaluation and achieves a 76-100% jailbreak discovery rate, outperforming baselines by up to 48 percentage points.

Zhiyuan Xu, Muhammad Firhard Roslan, Joseph Gardiner et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.