Public-facing health information on achalasia generated by large language models: a multidimensional comparative study
Abstract
Large language models (LLMs) are increasingly used by the public to obtain health information, but their ability to provide reliable information for achalasia remains unclear. We therefore compared six LLMs in answering public questions about achalasia, an uncommon esophageal motility disorder, focusing on safety, accuracy, empathy, information quality and reliability, and readability. In this cross-sectional comparative study, 40 patient-oriented questions about achalasia were submitted once to each of six LLMs (ChatGPT 5.5, Claude Opus 4.7, DeepSeek V4 Pro, Gemini 3.1 Pro Thinking, Grok-4.3, and Qwen3-Max), between May 7 and May 12, 2026. Three gastroenterologists independently assessed 240 responses for safety, accuracy, empathy, information quality, and readability using DISCERN, EQIP, the JAMA benchmark criteria assessing authorship, attribution, disclosure, and currency, and the Global Quality Score (GQS). Readability was evaluated using six established readability indices. Overall, 27 of 240 responses (11.3%) were classified as potentially unsafe, with proportions ranging from 7.5 to 15.0% across models; the overall between-model difference in safety was not statistically significant. Accuracy differed significantly among models, although the effect size was small. Larger differences were observed in empathy, information reliability and quality, and readability. ChatGPT 5.5 and Gemini 3.1 Pro Thinking achieved higher empathy scores, whereas Claude Opus 4.7 showed greater reading difficulty. JAMA benchmark scores were low overall, indicating limited source transparency. Current LLMs provided generally accurate and often useful answers to public questions about achalasia. However, some differences were found in safety, empathy, transparency, information quality, and readability. Public-facing LLM responses to public questions about achalasia should be consistent with guidelines, clearly cite sources, communicate with empathy, and use plain language.