This thesis sets out to answer whether appending “Your reasoning steps will be monitored” changes the sycophantic tendencies of an LLM, and shows a significant decrease in sycophantic behaviour.
It is demonstrated that detection performance drops on unseen models and an initial approach is proposed to address this challenge, and it is shown that detecting PSRS is feasible from the response text alone, and detectors need to learn subtle PSRS patterns from the training data.
Bohan Jiang, Dawei Li, Yasin N. Silva et al.· 0 citations
Large language models are increasingly used as sources of advice and information, including in high-stakes settings, yet little is known about how they respond to user disagreement. We study how a model manages its epistemic authority, referring here to its claim to knowledge, competence, or the right to advise, once a...
Riyadh Alnasser, Y. Çetinkaya, Su-Min Zhao et al.· 0 citations
In recent years, media attention has focused on artificial intelligence, particularly on chatbot services and generative intelligence. ChatGPT, created by OpenAI, was one of the earliest online tools and rapidly gained popularity. Users are indeed exposed to a service with privacy notifications and conditions of use th...
Jacopo Bassetta, D. Perpetuini, Maria Teresa Giusti et al.· Information· 0 citations
Large language models (LLMs) may abandon correct positions when users push back, exhibiting a failure mode known as sycophancy. Existing evaluations typically use short, pre-specified conversations and may therefore miss failures that emerge under sustained, adaptive disagreement. We introduce SPINE, a benchmark in whi...
Lei Tang, Kangda Wei, Tian-Yu Jiang et al.· 0 citations
The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc interpretability of static models to mechanistic understanding and proactive control of gen...
Rahul Gupta, Abhinav Mohanty, Anaelia Ovalle et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.