Skip to content
Preprint

Localizing Safety Alignment: MLP Layers and Mid-Network Blocks Encode Refusal Behavior in Large Language Models

Aug 2026 · 0 citations · 29 references
Computer Science

TL;DR

The results suggest that safety alignment in current LLMs is both localized and interaction-sensitive, offering insight into alignment brittleness and potential avenues for targeted safety interventions.

Abstract

Safety alignment in large language models is often treated as a distributed property of the entire network, yet its practical brittleness suggests that refusal behavior may be concentrated in a smaller set of parameters. This work addresses where safety-aligned refusal is encoded by transplanting weights from aligned models into matched unaligned base models at multiple levels of granularity. Using two open-weight model pairs and four safety benchmarks, we conducted experiments to compare the effects of replacing attention weights, MLP weights, contiguous layer regions, and MLP blocks. Across both model families, refusal transfer is dominated by MLP weights: replacing MLP parameters recovers substantially more malicious-prompt refusal than replacing attention parameters, with gains of at least 2.7 times more across benchmarks. Within the MLP stack, refusal-relevant parameters exhibit a consistent mid-network concentration, as the block spanning layers 8-11 is selected first in all six greedy searches over model-dataset pairs. The results also show that the composition of safety-relevant components is non-additive: in five of six greedy trajectories, adding more aligned blocks can reduce refusal performance, and selective block subsets can outperform full MLP transplantation on malicious refusal, benign over-refusal, or both. Finally, greedy orders transferred to OR-Bench vary with the source benchmark used to derive them, indicating a benchmark-dependent precision-coverage trade-off. These results suggest that safety alignment in current LLMs is both localized and interaction-sensitive, offering insight into alignment brittleness and potential avenues for targeted safety interventions.

View source

Similar papers

#artificial intelligence Preprint Oct 2026

SafeEvo: Deciphering the Safety Alignment Mechanism and Evolution in Language Models

Safety interpretability advances the study of Large Language Model (LLM) alignment from behavioral constraints driven by data or algorithms towards a deeper understanding of internal mechanisms. However, existing works have focused primarily on safety-related representations, attention heads, or neurons after alignment...

Miao Yu, Hao-Hao Huang, Luiza S. B. Yuan et al. · 0 citations
#natural language process... Preprint Sep 2026

Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness

No post-training method achieves all three properties the authors would want from safe alignment at once: refusal that isn't concentrated in a few fragile components, safety gains that don't cost general capability, and safety behavior correctable through small, targeted edits.

Hoang Cuong Nguyen, M. Dras, Usman Naseem · 1 citation
Preprint Aug 2026

HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models

Experiments across three instruction-tuned models show that HiRoute achieves high safety rates across multiple safety benchmarks while preserving safe-response helpfulness, reducing over-refusal, and maintaining competitive performance on general-purpose tasks.

Fang-Zhou Chen, Shiji Zhao, Mengyan Wang et al. · 0 citations
#machine learning Preprint Sep 2026

Refusal Localizes, the Damage Relocates: Safety Layers Under Few-Sample Fine-Tuning

Fine-tuning adapts aligned large language models (LLMs) to downstream tasks, but a few dozen harmful examples can remove their refusal of harmful requests. Prior work localizes safety-related behavior to specific layers, directions, and tokens, suggesting targets for protection. We test whether successful localization...

Jungseob Lee, D. Lee, Sugyeong Eo et al. · 1 citation
#artificial intelligence Preprint Sep 2026

SEAL: Reinforcing Global Safety in Mixture-of-Experts through Shared Expert ALignment

Mixture-of-Experts (MoE) is a scaling architecture for large language models that activates only a small subset of expert modules per token, enabling massive parameter growth with nearly constant computation. Recent Hybrid MoE architecture adds \textit{shared experts} to capture consistently useful representations, fur...

Qingyu Meng, Yiwei Zha, Jia-Huan Pei et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.