Skip to content

Similar papers

Preprint Aug 2026

Measuring and Detecting Harmful AI Sycophancy

It is demonstrated that detection performance drops on unseen models and an initial approach is proposed to address this challenge, and it is shown that detecting PSRS is feasible from the response text alone, and detectors need to learn subtle PSRS patterns from the training data.

Bohan Jiang, Dawei Li, Yasin N. Silva et al. · 0 citations
#artificial intelligence Preprint Sep 2026

How AI Models Manage Epistemic Authority: A Taxonomy and Comparative Analysis of Responses to User Disagreement

Large language models are increasingly used as sources of advice and information, including in high-stakes settings, yet little is known about how they respond to user disagreement. We study how a model manages its epistemic authority, referring here to its claim to knowledge, competence, or the right to advise, once a...

Riyadh Alnasser, Y. Çetinkaya, Su-Min Zhao et al. · 0 citations
Open access Sep 2026

The Readability of Generative AI Policies: An Empirical Analysis of OpenAI’s Informational Documents

In recent years, media attention has focused on artificial intelligence, particularly on chatbot services and generative intelligence. ChatGPT, created by OpenAI, was one of the earliest online tools and rapidly gained popularity. Users are indeed exposed to a service with privacy notifications and conditions of use th...

Jacopo Bassetta, D. Perpetuini, Maria Teresa Giusti et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Measuring LLM Sycophancy under Sustained Multi-Turn Pressure

Large language models (LLMs) may abandon correct positions when users push back, exhibiting a failure mode known as sycophancy. Existing evaluations typically use short, pre-specified conversations and may therefore miss failures that emerge under sustained, adaptive disagreement. We introduce SPINE, a benchmark in whi...

Lei Tang, Kangda Wei, Tian-Yu Jiang et al. · 0 citations
Preprint Aug 2026

From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop

The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc interpretability of static models to mechanistic understanding and proactive control of gen...

Rahul Gupta, Abhinav Mohanty, Anaelia Ovalle et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.