Skip to content

Evaluating Large Language Models on Misconceptions in Multi-Turn Medical Conversations

Jul 2026 · arXiv.org · Vol abs/2607.12884 · 0 citations · 44 references
Computer Science

TL;DR

ThReadMed-QA is introduced, a multi-turn medical dialogue dataset of 2,437 patient-physician conversation threads comprising 8,204 question-answer pairs that enables systematic evaluation of whether models can detect and correct misconceptions under a multi-turn context.

Abstract

Patients seeking medical information often ask questions that embed incorrect assumptions or misconceptions. In such cases, safe medical communication requires not only answering the question, but identifying and correcting the underlying false belief. These interactions naturally unfold over multiple turns, a pattern now mirrored in interactions with LLMs. Yet current evaluation frameworks do not capture model behavior in these settings, where misconceptions can emerge, persist, or evolve over the course of a conversation. Whether LLMs can reliably correct such misconceptions over time remains largely unexamined. To study this, we introduce ThReadMed-QA, a multi-turn medical dialogue dataset of 2,437 patient-physician conversation threads comprising 8,204 question-answer pairs, derived from real patient interactions on AskDocs. This dataset enables systematic evaluation of whether models can detect and correct misconceptions under a multi-turn context. We evaluate five LLMs using a rubric-based LLM-as-a-Judge framework that scores responses based on their ability to identify and correct misconceptions. Our experiments reveal a consistent pattern: even frontier models that can address misconceptions in a single interaction degrade substantially over subsequent turns. GPT-5 and Claude-Haiku correct these false presuppositions around 85% on initial questions but drop to roughly 50% within two follow-ups. An oracle analysis replacing prior model outputs with physician responses shows that much of the degradation is driven by error propagation, while performance remains imperfect even under correct context. Even when models tend to correct misconceptions initially, their performance degrades substantially over later turns, leading to inconsistent and potentially unsafe guidance in patient-facing settings and highlighting the need for evaluation frameworks that capture multi-turn behavior.

View source

Similar papers

Review

The Application of Large Language Models in Medical Question‑Answering Robots

This review explores the transformative impact of Large Language Models on medical question-answering robots and addresses critical challenges faced by LLM-driven systems, such as limited interpretability, restricted applicability across medical domains, and serious concerns about patient privacy.

Zhao-Xu Wang · 0 citations
Review

Let Me Explain!

Text content is the dominant factor in annotation decisions, far outweighing annotator demographics, and that content-focused SHAP explanations are more effective than demographic persona prompting for guiding LLM annotations, showing that explainability methods can improve both the reliability and the transparency of...

Hadi Mohammadi · 0 citations
#natural language process... Preprint Sep 2026

Evaluating Sycophancy in Chinese Large Language Models on Factual Questions Derived from Online Search Queries

This analysis covers 364,941 responses from three frontier Chinese-based LLMs to 12,165 factual questions derived from real-world Chinese search queries, and evaluates the models with and without reasoning across baseline, belief-conditioned, and anti-sycophancy prompting, tracing matched shifts among correct, incorrec...

Geng Liu, Feng Li, Meng-Xiao Zhu et al. · 0 citations
#natural language process... Preprint Aug 2026

WildSEEK: Evaluating Language Models for Information-Seeking

This work introduces WildSEEK, a manually annotated dataset of 3k information-seeking queries from real user interactions, and an evaluation framework for LLM-generated responses, and finds that over a third of information-seeking queries are high-risk and more often analytical.

Tanise Ceron, Joachim Baumann, Elisa Bassignana et al. · 0 citations
#natural language process... Preprint Sep 2026

Chronologic: Measuring Language Models'Ability to Represent the Past

Language models are appealing tools for research on the past. But to trust the evidence a model provides, researchers need to know whether its responses fit the period represented. Validation is challenging, because this is not a task living people ordinarily perform, and because many questions have multiple correct an...

Ted Underwood, Zi-Liang Qiu, Sarah Griebel et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.