The ASI is introduced, an Integrated Gradients-based token attribution method, which measures the degree to which a model's decision is driven by authority-related text, and a attribution-guided contrastive activation steering method is proposed to mitigate LLM sycophancy.
Abstract
Sycophancy refers to the tendency for large language models (LLMs) to match user beliefs at the cost of factual correctness, thereby undermining model reliability. Prior work on evaluating sycophancy in LLMs aims to assess whether a model's output matches an authority's claim, but cannot reveal which part of the prompt drives this sycophantic behavior. To bridge this gap, we investigate the relationship of sycophantic responses with an authority's credentials, their assertive claim, and the problem statement. We introduce the Authority Share Index (ASI), an Integrated Gradients-based token attribution method, which measures the degree to which a model's decision is driven by authority-related text. Through extensive experiments across five models and 30 test configurations, we find that sycophantic responses consistently direct more attention toward authority tokens than resistant ones. Moreover, our token attribution method reveals that for the sycophantic cases, the claim asserted by the authority receives more attention than the authority's credentials. Building on these findings, we propose attribution-guided contrastive activation steering to mitigate LLM sycophancy. Our method constructs a steering vector from high-attribution tokens of sycophantic and resistant responses, selectively pushing models toward resistance. This enables inference-time steering without retraining, lowering sycophancy from 96% to 25% in the strongest case. Together, our results show that token-level attribution can both explain what drives sycophancy and directly inform a practical intervention.
This thesis sets out to answer whether appending “Your reasoning steps will be monitored” changes the sycophantic tendencies of an LLM, and shows a significant decrease in sycophantic behaviour.
Victor Ruesink, Dr. Tom Kouwenhoven, C. Lennard et al.· 0 citations
This thesis presents a unified comparative analysis evaluating the robustness of three open-weights instruction-tuned models against a series of adversarial probing strategies spanning social, conversational, and analytical pressure, revealing that modern alignment strategies such as Reinforcement Learning from Human F...
Antia Alonso Cancela, Tom Kouwenhoven, Michiel van der Meer· 0 citations
It is demonstrated that detection performance drops on unseen models and an initial approach is proposed to address this challenge, and it is shown that detecting PSRS is feasible from the response text alone, and detectors need to learn subtle PSRS patterns from the training data.
Bohan Jiang, Dawei Li, Yasin N. Silva et al.· 0 citations
The findings suggest that the teacher models used to generate preference data can interact with alignment training objectives in unexpected ways, generalizing to undesirable and potentially harmful behaviors like sycophantic agreement.
C. Blank, Zhuo-Fan Ying, Christopher Potts et al.· 0 citations
Empirical evaluations show that two LRP-guided decoding methods that steer generation toward trajectories more relevant to cue tokens promote safer responses, supporting cue-token attribution's role in compliance failures.
Or Biton, Tomer Krichli, Itai Allouche et al.· 1 citation
A shared contrastive signal is introduced that identifies where sycophancy lives across the MoE hierarchy and drives interventions that act only where the behavior is present, demonstrating that sycophancy resides in identifiable computational subcircuits and can be selectively steered while maintaining a favorable rem...
Kareem Hassani, Chaymaa Abbas, Lama Mawlawi et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.