Skip to content
Open access

Evaluating large language models using the Type 2 Diabetes Health Education guideline: a comparative analysis of ChatGPT-4.1, Claude-4.0, DeepSeek-V3, and ERNIE Bot 4.5 Turbo

Aug 2026 · Frontiers in Public Health · Vol 14 · 0 citations · 44 references
Medicine

TL;DR

Although the four LLMs generally provide accurate and pertinent information regarding type 2 diabetes, enduring limits in actionability and inconsistencies among models in content completeness and understandability restrict their effective use in diabetic patient education.

Abstract

Objective To systematically evaluate the quality and readability of health information generated by four large language models (LLMs) in response to inquiries regarding type 2 diabetes mellitus (T2DM), using an authoritative Chinese clinical guideline as the reference standard. Methods A total of 124 standardized questions were extracted from the Chinese Type 2 Diabetes Popular Science Guidelines. Six endocrinologists and diabetes specialists conducted independent, blind evaluations using the CLEAR tool (Completeness, Lack of false Information, Evidence, Appropriateness, Relevance) and PEMAT-P (Patient Education Materials Assessment Tool for Printable materials). Response characteristics were also recorded. Between-model differences were tested using the Kruskal–Wallis H test with Bonferroni pairwise comparisons. Results All four models achieved total CLEAR scores within the “very good” range (19–25), with no significant differences seen between models (χ2 = 1.985, p = 0.576). No significant differences were observed in the dimensions of Lack of false information (χ2 = 7.644, p = 0.054), Evidence (χ2 = 2.309, p = 0.511), and Relevance (χ2 = 7.516, p = 0.057). However, significant differences emerged in Completeness (χ2 = 47.661, p < 0.001) and Appropriateness (χ2 = 88.360, p < 0.001). Claude-4.0 received the lowest score in Completeness (median 4.00, IQR 3.00–5.00) but achieved the highest ranking in Appropriateness (median 4.00, IQR 4.00–5.00). On the PEMAT-P, understandability differed significantly across models (χ2 = 159.120, p < 0.001), yet all models surpassed the 70% threshold, with ChatGPT-4.1 highest (median 91.91%, IQR 91.91–100.00%). However, despite significant differences among the various models (χ2 = 354.023, p < 0.001), only ERNIE Bot 4.5 Turbo (median 75.00%, IQR75.00–75.00%) surpassed the 70% threshold, with no single model demonstrating consistent superiority across all dimensions. Conclusion Although the four LLMs generally provide accurate and pertinent information regarding type 2 diabetes, enduring limits in actionability and inconsistencies among models in content completeness and understandability restrict their effective use in diabetic patient education. Future development should prioritize stronger step-by-step behavioral guidance and differentiated, scenario-specific model deployment to enhance their value in patient-facing diabetes self-management support.

Read PDF

Similar papers

Aug 2026

Accuracy and Safety of Language Model-Generated Breastfeeding Counseling Responses: An Expert-Based Comparative Evaluation of ChatGPT-3.5, ChatGPT-4, and BreastfeedGPT

Aim: This study aims to compare the responses generated by ChatGPT-3.5, ChatGPT-4, and BreastfeedGPT, a customized domain-specific GPT, to frequently asked questions related to breastfeeding counseling. Design and Methods: Ten breastfeeding-related questions were selected based on international guidelines and expert va...

Seda Serhatlıoğlu, Yeşim Yeşil · 0 citations
Open access Sep 2026

A Multidimensional Evaluation of Large Language Model Responses to the 2024 ESC Hypertension Guidelines: A Comparative Study

Large language models (LLMs), including ChatGPT, Gemini, and DeepSeek, are increasingly used in medicine; however, their performance across multiple clinically relevant domains remains incompletely understood. This cross-sectional comparative study evaluated 225 responses generated by three LLMs to 75 clinical question...

F. Böyük, Aysun Karahan Gün, İsmail Polat Canbolat et al. · 0 citations
Sep 2026

Evaluation of the accuracy and reproducibility of large language models (ChatGPT, DeepSeek, Gemini) in responding to patient-centered lipedema questions.

BackgroundLipedema is a frequently misdiagnosed chronic condition that significantly impacts patients' quality of life. As artificial intelligence (AI)-based large language models (LLMs) become increasingly integrated into healthcare communication, their accuracy and consistency in providing patient-centered informatio...

Rabia Sanır, E. Türkmen, E. Giray et al. · 0 citations
Open access Sep 2026

Performance of large language models in answering public questions about nutrition in cirrhosis: a comparative study.

BACKGROUND large language models (LLMs) are increasingly used for public health information, but their performance in nutrition advice for cirrhosis remains uncertain. We investigated four LLMs in answering public questions about cirrhosis nutrition across safety, accuracy, empathy, information reliability and quality,...

Jun-Zheng Li, Ying-Jie Wu, Man Yang et al. · 0 citations
Open access Sep 2026

Evaluating large language models for myocardial infarction public health education: a comparative study on information quality, transparency and readability

Objective Myocardial infarction (MI) is an acute, life-threatening cardiovascular disease, and high-quality, accessible public health education is vital for emergency management. This study systematically evaluates the quality, transparency, clinical accuracy, patient safety, and readability of information generated by...

Tai-Long Lv, Wen-Kai Bao, Shu-Di Li et al. · 0 citations
Open access Aug 2026

Mapping Gaps and Improvement Targets in Large Language Model-Generated Melanoma Patient Education in a Non-English Setting

How well large language models (LLM) handle Turkish melanoma patient education varies widely from one model to the next, and findings suggest that LLM-generated Turkish melanoma materials may be useful as preliminary educational drafts.

Nıyazı Çetın, A. Atılan · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.