Skip to content
Preprint

NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation

Aug 2026 · 0 citations · 78 references
Computer Science

TL;DR

NeuronFuzz is presented, a white-box fuzzing framework that exploits internal safety neurons as continuous execution feedback for LLM safety evaluation and achieves a 76-100% jailbreak discovery rate, outperforming baselines by up to 48 percentage points.

Abstract

Safety evaluation is critical for assessing whether aligned Large Language Models (LLMs) remain robust against jailbreak attacks. Existing automated testing methods, however, largely rely on response-level feedback: each candidate prompt typically requires generating a target-model response to evaluate its attack effectiveness. This process is expensive and, more importantly, provides only sparse guidance on strongly aligned models, where most candidates are rejected with the same failure outcome. This paper presents NeuronFuzz, a white-box fuzzing framework that exploits internal safety neurons as continuous execution feedback for LLM safety evaluation. A SafetyOracle converts safety-neuron activations into a continuous safety alarm score that serves as feedback for fuzzing and can be obtained during prefill, eliminating response generation from the fuzzing loop. To construct the SafetyOracle, NeuronFuzz uses template-invariant harmful and benign inputs and stability-aware selection to identify a compact set of safety neurons whose activations capture harmful-intent recognition. Moreover, since the safety alarm score is differentiable, NeuronFuzz uses its gradients to identify safety-sensitive template positions and a masked language model to generate fluent, context-compatible mutations while preserving original harmful payload and avoiding additional optimization variables. We evaluate NeuronFuzz across 21 text and multimodal models. Across five white-box source models, it achieves a 76-100% jailbreak discovery rate, outperforming baselines by up to 48 percentage points. Its optimized templates further transfer zero-shot to open-weight and six proprietary target models, achieving average ASR and top-5 ensemble ASR (EASR) of 69.6%/92.6% and 44.1%/60.0%, respectively.

View source

Similar papers

Jul 2026

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks

Mask2Shield (M2S), a masked-forward alignment method that trains a model under this functional pruning procedure, reduces successful recomputed pruning attacks from 80--279 to 1--44 out of 313 prompts while generally preserving four capability benchmarks.

Jincheng Ying, Ming-Hui Xu, Yinhao Xiao et al. · 0 citations
Preprint Aug 2026

NeuronGuard: Robust LLM Safety Alignment via Ablation-Aware Safety Signal Redistribution

A fine-tuning-stage defense that simultaneously hardens LLMs against both attack classes by redistributing safety signals across a broader set of neurons, and provides a formal guarantee that NeuronGuard strictly reduces the attack success rate (ASR) upper bound.

Anjun Gao, Yueyang Quan, Yu Xia et al. · 2 citations
Preprint Aug 2026

When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs

SkillSafe-Bench is introduced, a controlled benchmark that scores skill-merged models on static refusal, adaptive jailbreak robustness, and capability retention under a conservative two-judge AND rule, and the static effect of merging is base-conditional.

Yu Ma, Hongli Shi, Jing Li et al. · 0 citations
Preprint Aug 2026

Tripwire: Triggering Aligned Refusal via Statistically Certified Safety Neurons

Neuron- and path-level interventions offer the finest-grained route to defending large language models (LLMs) against jailbreak attacks, yet existing methods fall short of this promise, i.e., they often compromise model utility significantly. Specifically, one line of work suppresses toxic neurons to erase harmful sema...

Wei Zhao, Zhe Li, Pei-Xin Zhang et al. · 0 citations
Preprint Aug 2026

Circuit Discovery Helps Detect LLM Jailbreaking: A Mechanistic Interpretability Study

A mechanistic analysis of the jailbreaking behavior in a large-scale, safety-aligned LLM, focusing on LLaMA-2-7B-chat-hf identifies computational circuits responsible for generating affirmative responses to jailbreak prompts and uncovers key attention heads and MLP pathways that mediate adversarial prompt exploitation.

Paria Mehrbod, Boris Knyazev, Guy Wolf et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.