Skip to content
Review Open access

A comparative evaluation of generative AI chatbots for patient-oriented advice on painful diabetic peripheral neuropathy: safety, accuracy, guideline concordance, actionability, and readability

Sep 2026 · Frontiers in Public Health · Vol 14 · 0 citations · 36 references
Medicine

Abstract

Background Patients with painful diabetic peripheral neuropathy (PDPN) increasingly use generative artificial intelligence chatbots for information on symptoms, treatment, foot care, and when to seek professional help. Their usefulness depends on safety, accuracy, guideline concordance, actionability, and readability. Objective To compare five publicly accessible generative AI chatbots in answering standardized patient-oriented questions about PDPN. Methods: Sixty standardized English-language questions covering eight clinical domains were submitted once to ChatGPT, Gemini, Microsoft Copilot, DeepSeek, and Doubao in separate single-turn conversations, yielding 300 responses. Five reviewers independently assessed safety, accuracy, guideline concordance, and actionability using predefined criteria. Guideline concordance was scored against six mapped elements per question and converted to a percentage. Actionability was assessed using seven binary criteria with prespecified question-level applicability. Readability was evaluated using six established indices. Paired comparisons used Cochran’s Q test for safety and Friedman tests for non-binary outcomes, followed by multiplicity-adjusted pairwise analyses. Results All 300 responses were analyzed. Inter-rater agreement was high for safety (Fleiss’ κ = 0.874), accuracy [ICC (2,1) = 0.881], guideline concordance [ICC (2,1) = 0.874], and actionability [ICC (2,1) = 0.877]. Twenty-eight responses (9.3%) were classified as unsafe or potentially unsafe. Unsafe-response rates ranged from 5.0 to 15.0%, with no detected overall between-model difference (Cochran’s Q = 4.462, p = 0.347). Accuracy, guideline concordance, and actionability differed across models (all p < 0.001; Kendall’s W = 0.683, 0.730, and 0.556, respectively). ChatGPT generally achieved higher content-related scores, whereas Doubao scored lower. All readability indices also differed across models (all p < 0.001), with ChatGPT and Doubao producing less complex text and DeepSeek showing greater reading difficulty. Conclusion The five chatbots showed distinct performance patterns across content quality and readability. Although no overall safety difference was detected, every system generated at least one response with a plausible pathway to inappropriate self-management, delayed assessment, medication or product misuse, or preventable injury. Chatbots may support general patient education, but medication decisions, foot-risk assessment, and urgent-care triage require professional verification.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.