Skip to content

DeShortcut-Align: Decoupling Spurious Shortcuts for Robust Safety Alignment in Large Reasoning Models

Sep 2026 · 0 citations
Computer Science

TL;DR

DeShortcut-Align is proposed, a shortcut-decoupling alignment framework that reduces dependence on superficial cues that significantly improves robustness against template-stripping bypass attacks, substantially reduces over-refusal, and better preserves general-purpose reasoning capabilities, thereby mitigating the alignment tax commonly observed in safety training.

Abstract

Safety alignment of large reasoning models (LRMs) via supervised fine-tuning (SFT) and reinforcement learning (RL) often yields near-perfect safety scores, yet this apparent success comes at the cost of severe over-refusal and degraded general capabilities. Through systematic empirical analysis, we find that these failures are closely associated with the learning of spurious shortcuts rather than robust intent-sensitive safety evaluation. Specifically, we identify two dominant shortcuts: formatting shortcuts, where refusal behaviors are overly bound to structural prompt templates that frequently appear in safety alignment corpora; and lexical shortcuts, where sensitive keywords reflexively trigger refusals on benign queries. To mitigate reliance on these shortcuts, we propose DeShortcut-Align, a shortcut-decoupling alignment framework that reduces dependence on superficial cues. DeShortcut-Align operates across three coordinated stages: (1) Refusal Sensitivity Attribution, which masks input tokens to quantify their impact on the final refusal response distribution; (2) Attribution-Guided Contrastive Augmentation, which constructs benign contrastive samples using high-sensitivity tokens to mitigate lexical shortcuts; and (3) Counterfactual Consistency Regularization, which constructs template-ablated states via attention blinding to enforce decision consistency across SFT and RL, mitigating formatting shortcut dependence. Experiments on 7B and 14B models demonstrate that DeShortcut-Align significantly improves robustness against template-stripping bypass attacks (reducing performance drops by up to 72%), substantially reduces over-refusal by over 58%, and better preserves general-purpose reasoning capabilities, thereby mitigating the alignment tax commonly observed in safety training.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

First Token Matters: Understanding Safety Collapse in Large Reasoning Models

It is found that the refusal-related signal of LRMs drops sharply at the first generated token under harmful queries, which is associated with unsafe response generation, and this work proposes SafeToken, a lightweight inference-time intervention that injects a learned continuous safety anchor precisely at reasoning on...

Yi-Zheng Yang, Hai-Ning Yu, Yue-Chen Wang et al. · 0 citations
#natural language process... Preprint Sep 2026

Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness

No post-training method achieves all three properties the authors would want from safe alignment at once: refusal that isn't concentrated in a few fragile components, safety gains that don't cost general capability, and safety behavior correctable through small, targeted edits.

Hoang Cuong Nguyen, M. Dras, Usman Naseem · 1 citation
Preprint Aug 2026

HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models

Experiments across three instruction-tuned models show that HiRoute achieves high safety rates across multiple safety benchmarks while preserving safe-response helpfulness, reducing over-refusal, and maintaining competitive performance on general-purpose tasks.

Fang-Zhou Chen, Shiji Zhao, Mengyan Wang et al. · 0 citations
Book Open access Oct 2026

Abstract Concept Grounding through Safety Alignment in Multimodal Large Language Models

Abstract concepts such as harmfulness, legality, privacy, and fraud present a major challenge for multimodal large language models (MLLMs), as they require reasoning across visual, linguistic, and contextual information rather than direct perception. This paper investigates three questions: whether supervised alignment...

Ying-Xu Wang, Oliver Lemon · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.