It is demonstrated that detection performance drops on unseen models and an initial approach is proposed to address this challenge, and it is shown that detecting PSRS is feasible from the response text alone, and detectors need to learn subtle PSRS patterns from the training data.
Abstract
Sycophantic responses are becoming pervasive in large language models (LLMs), and prior work has pointed out that some of them could be harmful. This paper focuses on one harmful sycophancy: preference-induced stance reversal sycophancy (PSRS), where a model reverses an initial stance merely to align with a user's stated preference. While existing research mainly measures how sycophantic a model is, we go further and ask whether PSRS can also be detected automatically from a single response. To investigate this at scale, we introduce CAP (Contrastive Anchor Probing), a framework for collecting labeled PSRS data. Applying CAP to 17 open- and closed-source LLMs, we collect 290,460 labeled responses across 12 everyday-advice domains. We organize our study around three research questions. (1) How often does PSRS occur? (2) How well can it be detected? (3) How does detection generalize to unseen models? We first reveal that PSRS rates range from 5% to 56% across LLMs, with more capable models being less sycophantic. Next, we show that detecting PSRS is feasible from the response text alone, and detectors need to learn subtle PSRS patterns from the training data. Because new LLMs appear rapidly, detectors inevitably encounter unseen models, making cross-model generalization an important framework goal. We demonstrate that detection performance drops on unseen models and propose an initial approach to address this challenge. We will release our dataset and code to support future research.
Large language models (LLMs) may abandon correct positions when users push back, exhibiting a failure mode known as sycophancy. Existing evaluations typically use short, pre-specified conversations and may therefore miss failures that emerge under sustained, adaptive disagreement. We introduce SPINE, a benchmark in whi...
Lei Tang, Kangda Wei, Tian-Yu Jiang et al.· 0 citations
This thesis sets out to answer whether appending “Your reasoning steps will be monitored” changes the sycophantic tendencies of an LLM, and shows a significant decrease in sycophantic behaviour.
Victor Ruesink, Dr. Tom Kouwenhoven, C. Lennard et al.· 0 citations
Large language models often exhibit sycophancy, revising their answers to align with users when users push back. Such answer flips, however, can arise from different causes. One possibility is that the model simply aligns with the user's feedback in order to satisfy them. Another is that the feedback genuinely contains...
Huanhuan Ma, Henry Peng Zou, Cheng-Ze Li et al.· 0 citations
This thesis presents a unified comparative analysis evaluating the robustness of three open-weights instruction-tuned models against a series of adversarial probing strategies spanning social, conversational, and analytical pressure, revealing that modern alignment strategies such as Reinforcement Learning from Human F...
Antia Alonso Cancela, Tom Kouwenhoven, Michiel van der Meer· 0 citations
Reliable refusal of harmful requests is essential to the safe deployment of language models. Because excessive eagerness to please users may undermine existing refusal capabilities, reducing sycophancy offers a potential route to stronger refusal beyond the harmful scenarios covered by safety training. We investigate t...
Large language models (LLMs) that have been fine-tuned using Reinforcement Learning from Human Feedback (RLHF) are likely to conform to the user, even when the user is wrong. This is one of the more intransigent side effects of current alignment methods and is called sycophancy. This review traces the origin of sycopha...
Venkata Phanindra Gollapalli, Sai M. Dasari, Shailesh Kadam et al.· 2026 International Conferenc...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.