ThReadMed-QA is introduced, a multi-turn medical dialogue dataset of 2,437 patient-physician conversation threads comprising 8,204 question-answer pairs that enables systematic evaluation of whether models can detect and correct misconceptions under a multi-turn context.
Abstract
Patients seeking medical information often ask questions that embed incorrect assumptions or misconceptions. In such cases, safe medical communication requires not only answering the question, but identifying and correcting the underlying false belief. These interactions naturally unfold over multiple turns, a pattern now mirrored in interactions with LLMs. Yet current evaluation frameworks do not capture model behavior in these settings, where misconceptions can emerge, persist, or evolve over the course of a conversation. Whether LLMs can reliably correct such misconceptions over time remains largely unexamined. To study this, we introduce ThReadMed-QA, a multi-turn medical dialogue dataset of 2,437 patient-physician conversation threads comprising 8,204 question-answer pairs, derived from real patient interactions on AskDocs. This dataset enables systematic evaluation of whether models can detect and correct misconceptions under a multi-turn context. We evaluate five LLMs using a rubric-based LLM-as-a-Judge framework that scores responses based on their ability to identify and correct misconceptions. Our experiments reveal a consistent pattern: even frontier models that can address misconceptions in a single interaction degrade substantially over subsequent turns. GPT-5 and Claude-Haiku correct these false presuppositions around 85% on initial questions but drop to roughly 50% within two follow-ups. An oracle analysis replacing prior model outputs with physician responses shows that much of the degradation is driven by error propagation, while performance remains imperfect even under correct context. Even when models tend to correct misconceptions initially, their performance degrades substantially over later turns, leading to inconsistent and potentially unsafe guidance in patient-facing settings and highlighting the need for evaluation frameworks that capture multi-turn behavior.
This review explores the transformative impact of Large Language Models on medical question-answering robots and addresses critical challenges faced by LLM-driven systems, such as limited interpretability, restricted applicability across medical domains, and serious concerns about patient privacy.
Text content is the dominant factor in annotation decisions, far outweighing annotator demographics, and that content-focused SHAP explanations are more effective than demographic persona prompting for guiding LLM annotations, showing that explainability methods can improve both the reliability and the transparency of...
The results show that medical sycophancy depends as much on how a model is challenged and evaluated as on which model is tested, and fabricated evidence has opposite effects across interaction structures.
Kaike Ping, Buse Çarik, Caleb Wohn et al.· 2 citations
This analysis covers 364,941 responses from three frontier Chinese-based LLMs to 12,165 factual questions derived from real-world Chinese search queries, and evaluates the models with and without reasoning across baseline, belief-conditioned, and anti-sycophancy prompting, tracing matched shifts among correct, incorrec...
Geng Liu, Feng Li, Meng-Xiao Zhu et al.· 0 citations
This work introduces WildSEEK, a manually annotated dataset of 3k information-seeking queries from real user interactions, and an evaluation framework for LLM-generated responses, and finds that over a third of information-seeking queries are high-risk and more often analytical.
Tanise Ceron, Joachim Baumann, Elisa Bassignana et al.· 0 citations
Language models are appealing tools for research on the past. But to trust the evidence a model provides, researchers need to know whether its responses fit the period represented. Validation is challenging, because this is not a task living people ordinarily perform, and because many questions have multiple correct an...
Ted Underwood, Zi-Liang Qiu, Sarah Griebel et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.