Preprint
Aug 2026
Measuring and Detecting Harmful AI Sycophancy
It is demonstrated that detection performance drops on unseen models and an initial approach is proposed to address this challenge, and it is shown that detecting PSRS is feasible from the response text alone, and detectors need to learn subtle PSRS patterns from the training data.
Bohan Jiang, Dawei Li, Yasin N. Silva et al.
· 0 citations