Mitigating Sycophancy and Alignment Failures in Large Language Models
Large language models (LLMs) that have been fine-tuned using Reinforcement Learning from Human Feedback (RLHF) are likely to conform to the user, even when the user is wrong. This is one of the more intransigent side effects of current alignment methods and is called sycophancy. This review traces the origin of sycopha...