Skip to content
Open access

Safety, accuracy, empathy, reliability, and readability of large language model chatbot responses to public-facing vegetarian and vegan nutrition advice questions: a cross-sectional comparative study

Aug 2026 · Frontiers in Public Health · Vol 14 · 0 citations · 55 references
Medicine

Abstract

Background Publicly accessible large language model (LLM) chatbots are increasingly used to seek nutrition advice. Although vegetarian and vegan nutrition advice is often perceived as low risk, it may involve supplementation, vulnerable life stages, chronic disease, and symptoms that require clinical assessment. Objective To compare the safety, accuracy, empathy, reliability, information quality, and readability of five publicly accessible LLM chatbot products in responding to public-facing vegetarian and vegan nutrition advice questions. Methods In this cross-sectional comparative study, 58 predefined public-facing vegetarian and vegan nutrition advice questions were submitted once in English to ChatGPT, Gemini, Microsoft Copilot, DeepSeek, and Doubao through their official web interfaces. The final dataset included 290 chatbot responses. Five blinded raters with clinical nutrition training evaluated safety, accuracy, and empathy using predefined reference-answer anchors, and assessed reliability and information quality using four established instruments: the DISCERN instrument, Ensuring Quality Information for Patients (EQIP), Journal of the American Medical Association (JAMA) benchmark criteria, and Global Quality Score (GQS). Readability was assessed using six formula-based indices: the Automated Readability Index (ARI), Coleman–Liau Index (CL), Flesch–Kincaid Grade Level (FKGL), Gunning Fog Index (GFI), Simple Measure of Gobbledygook (SMOG), and Flesch Reading Ease Score (FRES). Model comparisons were paired by question. Results Inter-rater agreement was high for all human-rated outcomes. Potentially harmful responses occurred in all models, although adjusted pairwise safety comparisons were not statistically significant. Recurring harm patterns involved vitamin B12 source reliability, uncontrolled iodine or selenium intake, vulnerable life stages, chronic disease, symptom triage, and poorly traceable or overconfident statements. Accuracy, empathy, reliability, information quality, and readability differed significantly across models. ChatGPT, Copilot, and DeepSeek showed higher accuracy; ChatGPT showed higher empathy; Copilot performed best on DISCERN and EQIP; and Gemini produced the most readable responses. Conclusion Publicly accessible LLM chatbots can provide useful general information on vegetarian and vegan nutrition, but their performance is uneven and safety limitations remain. Chatbot-generated advice should be interpreted cautiously, especially for supplementation, vulnerable groups, chronic disease, and symptoms requiring clinical assessment. Future systems should improve safety guidance, referral cues, source traceability, and plain-language communication.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.