Skip to content

Evaluating Sycophancy in Chinese Large Language Models on Factual Questions Derived from Online Search Queries

Sep 2026 · 0 citations · 8 references
Computer Science

TL;DR

This analysis covers 364,941 responses from three frontier Chinese-based LLMs to 12,165 factual questions derived from real-world Chinese search queries, and evaluates the models with and without reasoning across baseline, belief-conditioned, and anti-sycophancy prompting, tracing matched shifts among correct, incorrect, and uncertain responses.

Abstract

As large language models increasingly mediate information access, factually accurate and independent answers are critical. However, these models can exhibit sycophancy by aligning their responses with users'stated beliefs even when those beliefs are incorrect, potentially presenting misinformation as independently verified and reinforcing users'confidence in false claims. Prior work leaves unresolved whether introducing user beliefs causes correct responses to become incorrect or uncertain, or causes uncertain responses to become belief-aligned incorrect answers. It also remains unclear whether anti-sycophancy interventions preserve or restore factual accuracy or merely shift responses toward uncertainty. We analyze factual sycophancy in Chinese-language information seeking using yes/no fact-checking questions. Our analysis covers 364,941 responses from three frontier Chinese-based LLMs (DeepSeek, Qwen, and Doubao) to 12,165 factual questions derived from real-world Chinese search queries. We evaluate the models with and without reasoning across baseline, belief-conditioned, and anti-sycophancy prompting, tracing matched shifts among correct, incorrect, and uncertain responses. Under incorrect user beliefs, we distinguish belief-aligned errors from losses of factual confidence, in which initially correct answers become uncertain. Patterns vary across models and reasoning settings: reasoning is not a consistent safeguard, and anti-sycophancy instructions can reduce incorrect agreement while increasing uncertainty. In Chinese-language factual question answering, avoiding agreement with false beliefs is therefore not equivalent to preserving factual accuracy, highlighting the value of transition-level evaluation. Such behavior may undermine the reliability of LLM-mediated information access by reinforcing misinformation or weakening users'confidence in factually correct answers.

View source

Similar papers

#natural language process... Preprint Sep 2026

Authority Bias in Language Models: Source Deference and User Agreement Are Not Interchangeable

Language models tend to agree with whatever a user asserts, and post-training increasingly targets this sycophancy so that models evaluate claims on their merits rather than deferring to the user. Yet the same models are far more compliant when a wrong answer is attributed to a verified source, which is how retrieval r...

Abhinav Kumar, Paras Chopra · 0 citations
#artificial intelligence Preprint Sep 2026

You Really Didn't Get That? Benchmarking Social Pragmatic Inference for Indirect and Playful Chinese Online Comments

A benchmark for evaluating whether LLMs can recover situated pragmatic meanings in Chinese online comments is introduced, and case analysis shows that models often recognize broad irony or playfulness while misidentifying the mechanism or interactional move.

Yin-Hong Shi, Jun-Jie Ma, Emma Jiren Wang et al. · 0 citations
2026

Beyond Literal Meaning: How LLMs Interpret Yemeni Proverbs

Results show that instruction-tuned models like GPT-4o and Gemini 1.5 Pro outperform smaller models in both automatic and human evaluations, and LLM-as-a-Judge evaluation correlates strongly with human assessment.

Nasser Thmer, Ali Allaith, Muhammad Shoaib · 0 citations
Preprint Aug 2026

SAGE: From Direct Answering to Evidence-Grounded Inference for Chinese Ancient Document Understanding

SAGE is an evidence-grounded multi-agent framework that reformulates Chinese ancient document understanding as evidence-grounded inference rather than direct answer generation, highlighting the importance of structured, evidence-grounded inference beyond model scaling.

Yu-Chuan Wu, Xuan Luo, Yinglian Zhu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Fair Fact-Checking: Closing the Cross-Lingual Gap in LLM Factual Judgement with RoSh

Misinformation on social media remains a critical problem, and more and more people settle it by asking a language model instead of a fact checker. Whether models judge such claims reliably is debated; whether they judge them equally well in every language people ask in has gone almost unasked. We test eight models fro...

Muhammad Ahmad, Fatemeh Seyedin, Adrian Weller et al. · 0 citations
#natural language process... Preprint Sep 2026

When Models Defer to Wrong Answers: A Robustness Audit of Source-Attributed Cues in Multiple-Choice QA

Language models often receive a question together with a claim about what another source answered. We audit whether such claims destabilize answers in multiple-choice question answering. For each item, we hold one wrong option fixed across misleading conditions and vary the cue template attached to it. We introduce \em...

Manikandan Ravikiran, Siddharth Vohra · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.