Evaluating Sycophancy in Chinese Large Language Models on Factual Questions Derived from Online Search Queries
This analysis covers 364,941 responses from three frontier Chinese-based LLMs to 12,165 factual questions derived from real-world Chinese search queries, and evaluates the models with and without reasoning across baseline, belief-conditioned, and anti-sycophancy prompting, tracing matched shifts among correct, incorrec...