Skip to content
Review Open access

Evaluation of generative AI-driven chatbots as sources of consumer health information on hand, foot, and mouth disease: a cross-sectional comparative study of safety, accuracy, information quality, readability, and empathy

Aug 2026 · Frontiers in Public Health · Vol 14 · 0 citations · 45 references
Medicine

TL;DR

Across a limited set of standardized English-language prompts, the five chatbots often generated coherent HFMD information, but occasional safety-relevant inaccuracies and readability barriers remained, suggesting these systems may support general information seeking but should not replace individualized professional advice.

Abstract

Background Generative artificial intelligence chatbots are increasingly used as sources of consumer health information. Although their performance has been examined in several medical conditions, evidence specific to hand, foot, and mouth disease (HFMD) remains limited. Objective To compare the safety, accuracy, empathy, information quality, and readability of HFMD-related responses generated by five publicly accessible chatbots. Methods In this exploratory cross-sectional study, 20 researcher-developed English-language prompts were constructed from authoritative public-health sources, Google Trends topic mapping, and caregiver-informed wording refinement. ChatGPT-4o, Gemini 2.5 Pro, Copilot, Doubao, and DeepSeek-V3.2-Exp were evaluated between April 2 and April 5, 2026. Five trained reviewers independently assessed responses against predefined reference standards using safety and accuracy criteria, an empathy scale, DISCERN, EQIP, the Global Quality Score, and JAMA benchmarks. Six formula-based readability indices were calculated. Paired comparisons used Friedman and Cochran Q tests, with prespecified post-hoc procedures and Benjamini-Hochberg correction. Results Unsafe-response rates ranged from 5.0 to 15.0%, with no statistically significant difference detected among chatbots (p = 0.797). No statistically significant inter-model differences were detected for accuracy, empathy, DISCERN, EQIP, JAMA, or Global Quality Score in the 20-prompt set; because the study was exploratory and was not powered to establish equivalence, these findings do not demonstrate comparable or interchangeable performance. All six readability indices differed significantly among chatbots (p = 0.030 to <0.001). ChatGPT and Doubao generally produced lower estimated grade-level complexity than Gemini and DeepSeek. Ten potentially unsafe or misleading responses were identified, mainly involving overgeneralization of EV71 vaccine protection, hand-hygiene qualification, and disinfection advice. Conclusion Across a limited set of standardized English-language prompts, the five chatbots often generated coherent HFMD information, but occasional safety-relevant inaccuracies and readability barriers remained. These systems may support general information seeking, but their responses require cautious interpretation and should not replace individualized professional advice.

Read PDF

Similar papers

#large language models Review Open access Sep 2026

A cross-sectional evaluation of large language model chatbot interfaces for patient-facing herpes zoster information: safety, information quality, and readability

Background Large language model (LLM) chatbots are increasingly used to obtain health information. However, fluent and clinically plausible responses may still contain safety-relevant omissions, inadequate source attribution and disclosure, or difficult-to-read text. Objective To evaluate the safety, accuracy, empathy,...

Da-Lian Liang, Cong Mai, Zheng-Kun Zhang et al. · 0 citations
Open access Sep 2026

Cross-sectional comparative evaluation of five large language model–driven chatbots on an expert-curated 44-question set of parent-facing questions about pediatric vitamin D deficiency

Background Parents and caregivers increasingly use generative artificial intelligence chatbots for child health information. Pediatric vitamin D deficiency is a clinically relevant topic because advice about supplementation, testing, rickets, high-risk children, toxicity, and emergency symptoms can influence caregiver...

Ping Shi, Tian Zhou, Qiao Nie et al. · 0 citations
Open access Sep 2026

Quality and Readability of Four AI Chatbots Answering Search-Derived Patient Questions About Primary Aldosteronism: A Cross-Sectional Comparative Study

Highlights What are the main findings? Instrument-based information-quality scores differed among the tested chatbot configurations, but these differences do not establish greater factual accuracy, guideline concordance, or clinical superiority; the full-set DISCERN comparison was limited because most questions were no...

Jian-Yu Chen, Zai-Han Zhao, Shang-Yu Han et al. · 0 citations
#small language model Open access Sep 2026

Safety, accuracy, empathy, information quality, and readability of publicly accessible LLM-based chatbots for traumatic brain injury and concussion questions: a cross-sectional comparative study

Background Large language model (LLM)-based chatbots are increasingly used by the public to obtain health information, but their performance in answering questions related to traumatic brain injury (TBI) and concussion remains unclear. This study evaluated five publicly accessible LLM-based chatbots across safety, accu...

Xin Zuo, Huan Zuo, Min Zhang et al. · 0 citations
Open access Aug 2026

Safety and quality of public chatbots for lung cancer prognostic information: a comparative evaluation

Public-facing chatbots may support general patient education but should not replace individualized clinician-led prognostic communication as public-facing chatbots differed substantially in safety, reliability, communication quality, and readability.

Yan-Ru Jiang, Qian-Yun Wang, Liang Zheng et al. · 0 citations
#generative ai Open access Sep 2026

A comparative cross-sectional evaluation of generative AI chatbots for patient-oriented bacterial vaginosis health advice: safety, accuracy, guideline concordance, empathy, and readability

Bacterial vaginosis (BV) concerns involve intimate symptoms, stigma, diagnostic uncertainty, medication use, pregnancy, sexual health, and self-care. Publicly accessible generative artificial intelligence chatbots offer immediate and potentially non-judgmental information, but the quality and safety of specific res...

Ling Miao, Li-Na Gu, Zhao-Le Gong et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.