Skip to content
Open access

Cross-sectional comparative evaluation of five large language model–driven chatbots on an expert-curated 44-question set of parent-facing questions about pediatric vitamin D deficiency

Sep 2026 · Frontiers in Pediatrics · Vol 14 · 0 citations · 45 references
Medicine

Abstract

Background Parents and caregivers increasingly use generative artificial intelligence chatbots for child health information. Pediatric vitamin D deficiency is a clinically relevant topic because advice about supplementation, testing, rickets, high-risk children, toxicity, and emergency symptoms can influence caregiver decisions. Objective To compare the safety, medical accuracy, empathy, information reliability, educational quality, transparency, global quality, and readability of five large language model-driven chatbots when answering questions about pediatric vitamin D deficiency. Methods This cross-sectional comparative study was reported with reference to CHART. An expert-curated 44-question set was selected from a 92-question candidate pool informed by search trends, caregiver-facing sources, clinical guidelines, and expert discussion. Each of the 44 questions was submitted once to each of five chatbot services—ChatGPT-5.5, Gemini 3.1 Pro, Qianwen 3.6-Plus, DeepSeek V4, and Doubao-Seed-2.0 Pro—using the same parent-oriented instruction. Responses were assessed for safety, accuracy, empathy, DISCERN, EQIP, JAMA criteria, GQS, and readability. Paired question-level differences were analyzed using Friedman tests, Kendall's W, and Cochran's Q; results were interpreted as a single-run, time-specific snapshot. Results Inter-rater agreement was good to excellent (Fleiss' kappa = 0.842 for safety; ICCs 0.846–0.914 for other subjective metrics). Twenty of 220 responses (9.1%) were unsafe; safety did not differ significantly across models (Cochran's Q = 3.704, df = 4, P = 0.448). Significant inter-model differences were observed in accuracy, empathy, DISCERN, EQIP, JAMA, GQS, and all readability indices. ChatGPT had the highest median accuracy [5.00 (4.00, 5.00)] and DISCERN score [69.60 (65.25, 72.40)]; DeepSeek and Doubao had the highest empathy scores [5.00 (4.80, 5.00)]. ChatGPT and Doubao shared the highest EQIP median (86.00), Doubao had the highest GQS [5.00 (4.00, 5.00)], and Gemini had the highest FRES. JAMA scores were low across models. Conclusions In this single-run evaluation, the five chatbot services showed domain-specific performance differences without a significant safety difference. Findings are exploratory rather than evidence of stable model superiority and support multidimensional evaluation with clinician involvement for high-risk or individualized pediatric advice.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.