Skip to content

Token-Level Diagnosis of Sycophancy in LLMs with Attribution-Guided Steering

Jul 2026 · arXiv.org · Vol abs/2607.28906 · 1 citation · 51 references
Computer Science

TL;DR

The ASI is introduced, an Integrated Gradients-based token attribution method, which measures the degree to which a model's decision is driven by authority-related text, and a attribution-guided contrastive activation steering method is proposed to mitigate LLM sycophancy.

Abstract

Sycophancy refers to the tendency for large language models (LLMs) to match user beliefs at the cost of factual correctness, thereby undermining model reliability. Prior work on evaluating sycophancy in LLMs aims to assess whether a model's output matches an authority's claim, but cannot reveal which part of the prompt drives this sycophantic behavior. To bridge this gap, we investigate the relationship of sycophantic responses with an authority's credentials, their assertive claim, and the problem statement. We introduce the Authority Share Index (ASI), an Integrated Gradients-based token attribution method, which measures the degree to which a model's decision is driven by authority-related text. Through extensive experiments across five models and 30 test configurations, we find that sycophantic responses consistently direct more attention toward authority tokens than resistant ones. Moreover, our token attribution method reveals that for the sycophantic cases, the claim asserted by the authority receives more attention than the authority's credentials. Building on these findings, we propose attribution-guided contrastive activation steering to mitigate LLM sycophancy. Our method constructs a steering vector from high-attribution tokens of sycophantic and resistant responses, selectively pushing models toward resistance. This enables inference-time steering without retraining, lowering sycophancy from 96% to 25% in the strongest case. Together, our results show that token-level attribution can both explain what drives sycophancy and directly inform a practical intervention.

View source

Similar papers

Review

The Impact of Stated Reasoning Monitoring on Sycophantic Response Tendencies in LLMs

This thesis sets out to answer whether appending “Your reasoning steps will be monitored” changes the sycophantic tendencies of an LLM, and shows a significant decrease in sycophantic behaviour.

Victor Ruesink, Dr. Tom Kouwenhoven, C. Lennard et al. · 0 citations

Analysis of Sycophancy Across Question-ing Styles A Comparative Study of Sycophancy Across Prompting Architectures and Large Language Models

This thesis presents a unified comparative analysis evaluating the robustness of three open-weights instruction-tuned models against a series of adversarial probing strategies spanning social, conversational, and analytical pressure, revealing that modern alignment strategies such as Reinforcement Learning from Human F...

Antia Alonso Cancela, Tom Kouwenhoven, Michiel van der Meer · 0 citations
Preprint Aug 2026

Measuring and Detecting Harmful AI Sycophancy

It is demonstrated that detection performance drops on unseen models and an initial approach is proposed to address this challenge, and it is shown that detecting PSRS is feasible from the response text alone, and detectors need to learn subtle PSRS patterns from the training data.

Bohan Jiang, Dawei Li, Yasin N. Silva et al. · 0 citations
Preprint Aug 2026

Hidden in the Request: Explaining Unethical LLM Compliance through Token Relevance

Empirical evaluations show that two LRP-guided decoding methods that steer generation toward trajectories more relevant to cue tokens promote safer responses, supporting cue-token attribution's role in compliance failures.

Or Biton, Tomer Krichli, Itai Allouche et al. · 1 citation
Preprint Aug 2026

THESIS-MoE: Trainable Hierarchical Extraction and SteerIng of Sycophancy in Mixture-of-Experts

A shared contrastive signal is introduced that identifies where sycophancy lives across the MoE hierarchy and drives interventions that act only where the behavior is present, demonstrating that sycophancy resides in identifiable computational subcircuits and can be selectively steered while maintaining a favorable rem...

Kareem Hassani, Chaymaa Abbas, Lama Mawlawi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.