Skip to content
Preprint

Hidden in the Request: Explaining Unethical LLM Compliance through Token Relevance

Aug 2026 · 1 citation · 31 references
Computer Science

TL;DR

Empirical evaluations show that two LRP-guided decoding methods that steer generation toward trajectories more relevant to cue tokens promote safer responses, supporting cue-token attribution's role in compliance failures.

Abstract

Although Large Language Models (LLMs) are aligned to optimize for both helpfulness and harmlessness, these dual objectives may conflict, inevitably leading to alignment failures. This work systematically investigates instances where LLMs fail to exhibit ethical behavior. To understand the underlying mechanics of these vulnerabilities, we introduce a probing methodology that presents unethical scenarios to LLMs in three distinct structural modalities: objective classification tasks, subjective first-person statements, and direct requests for assistance. We find that model performance degrades in the request-for-assistance-based form. Using Layer-wise Relevance Propagation (LRP), we trace this discrepancy to an attribution bias: the model places greater emphasis on benign task-framing tokens (e.g.,"Can you help me...") than on tokens signaling the underlying unethical behavior (e.g.,"without getting caught"), which we term cue-tokens. We hypothesize that this under-attribution contributes to harmful compliance. To test this, we introduce two LRP-guided decoding methods that steer generation toward trajectories more relevant to cue tokens. Empirical evaluations show that these interventions promote safer responses, supporting cue-token attribution's role in compliance failures.

View source

Similar papers

Preprint Aug 2026

Mood Matters: How Syntactic Sensitivity Undermines Safety Alignment

Large language models typically undergo post-training to align them with safety policies but there exist many sophisticated jailbreaks that sidestep established safeguards. For instance, prior work by Andriushchenko et al. (2025) has found that changing the grammatical tense from present to past can be enough to elicit...

Alina Klerings, Jannik Brinkmann, Heiner Stuckenschmidt et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models

Striking a balance between helpfulness and safety remains a fundamental challenge in aligning large language models. To achieve this balance, models should refuse harmful queries (e.g.,"How do I shoot someone?") while remaining responsive to benign inputs, even those superficially resembling harmful queries (e.g.,"Wher...

Minji Kim, Hyounghun Kim · 2 citations
Jul 2026

Adaptive Capitulation: A Structural Failure Mode of LLM Responses in Vulnerability Contexts

Large language models operating in emotionally sensitive contexts face a structural trilemma: when users in vulnerable states request information that may reinforce maladaptive attribution, current response architectures resolve the tension through protective restriction, uninflected facilitation, or unintegrated co-pr...

E. Lee · 0 citations
Jul 2026

Token-Level Diagnosis of Sycophancy in LLMs with Attribution-Guided Steering

The ASI is introduced, an Integrated Gradients-based token attribution method, which measures the degree to which a model's decision is driven by authority-related text, and a attribution-guided contrastive activation steering method is proposed to mitigate LLM sycophancy.

H. Nguyen, M. Kamruzzaman, Anshuman Chhabra et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.