Skip to content
Preprint

Mitigating LLM sycophancy with RL-based fine-tuning: Bayesian Truth Serum approach

Aug 2026 · 0 citations
Computer Science Engineering

TL;DR

A methodology for mitigating sycophancy that employs the Bayesian Truth Serum (BTS), a peer-prediction mechanism, as the reward in Group Relative Policy Optimization (GRPO) to fine-tune an LLM.

Abstract

Large language models (LLMs) frequently exhibit \emph{sycophancy}: they adapt their answers to a user's stated beliefs or preferences instead of reporting what they hold to be true, which lowers factual accuracy and can amplify misinformation. This paper proposes a methodology for mitigating sycophancy that employs the Bayesian Truth Serum (BTS), a peer-prediction mechanism, as the reward in Group Relative Policy Optimization (GRPO) to fine-tune an LLM. BTS pays an answer for being \emph{surprisingly common}, that is, more frequent among respondents than those respondents themselves predicted. We treat a group of responses from a model for one question as those respondents, so the reward is a function of the model's own outputs and fine-tuning needs neither labels nor preference annotations. We prove that in the large-group limit a sycophantic response earns strictly lower expected reward than an honest one. We also prove that if the entire group agrees in advance on a symmetric answering rule, it cannot earn a higher information score than under truthful reporting. On our true/false benchmark the reference model's answer-flip rate under user pressure decreases from 23% to 4%, and its accuracy under that pressure increases from 80% to 93%. Our reward outperforms SMART and is comparable to synthetic-data fine-tuning and to pinpoint tuning, all three of which train on labels. It spends considerably more compute in exchange, which makes it suitable when labeled data is scarce. Peer Truth Serum, which also pays a premium for a rare answer but elicits no prediction report, reproduces the effect. A peer-prediction reward computed inside a single GRPO group therefore reduces sycophancy without labels, and comparing mechanisms suggests that the premium paid for a rarer answer drives the effect.

View source

Similar papers

Conference Aug 2026

Mitigating Sycophancy and Alignment Failures in Large Language Models

Large language models (LLMs) that have been fine-tuned using Reinforcement Learning from Human Feedback (RLHF) are likely to conform to the user, even when the user is wrong. This is one of the more intransigent side effects of current alignment methods and is called sycophancy. This review traces the origin of sycopha...

Venkata Phanindra Gollapalli, Sai M. Dasari, Shailesh Kadam et al. · 0 citations
#artificial intelligence Preprint Sep 2026

OpenJev-RLCD: A Working RLCD Implementation

Decision models such as Jev answer questions with probabilities, which are only useful if they are calibrated. Open-source reproductions rely on supervised fine-tuning plus temperature scaling, while reinforcement learning from verifiable rewards (RLVR) makes reasoning models overconfident. We present a working impleme...

Zhi-Min Gao, Pi-Chao Wang · 0 citations
#machine learning Preprint Sep 2026

Distortion of AI Alignment Revisited: RLHF is a Decent Utilitarian Aligner

While Reinforcement Learning from Human Feedback (RLHF) is the standard paradigm for aligning large language models with human preferences, its effectiveness in pluralistic settings has been called into question. Notably, recent work by G\"olz et al. (2025) demonstrated that the \textit{distortion} -- defined as the mu...

Kazusato Oko, Annie Ulichney, Nika Haghtalab et al. · 3 citations · ⚡1
#artificial intelligence Preprint Sep 2026

Spurious Advantage Hidden in GRPO

SignBALANCE is proposed, whose magnitude is composition-free: it keeps the verifier sign, uses a global scale, and restores zero-mean balance via a stop-gradient per-class rescaling.

Jiamian Wang, Samyadeep Basu, Koustava Goswami et al. · 0 citations
#machine learning Preprint Aug 2026

An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning

This work compares four reward designs that span lexical suppression, anti-refusal shaping, rubric-based broad answering, and an explicit refusal contrast, and shows that optimization success is not equivalent to behavioral unlearning: RWKU forget scores, held-out completion audits, terminal training-rollout audits, an...

Rubén Balbastre, J. Orduña, M. Perez · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.