Aug 2026· Frontiers in Medicine· Vol 13· 0 citations· 17 references
Medicine
TL;DR
While most responses aligned with the IDF-DAR guideline, some harmful recommendations persisted, and response consistency fluctuated over time, the results suggest that AI chatbots should serve as a supplementary resource rather than a substitute for professional medical advice.
Abstract
Background Patients with diabetes increasingly consult artificial intelligence (AI) chatbots for medical advice, including guidance on antidiabetic medication management during Ramadan fasting, because AI can simplify and summarize long, complex guidelines. Also, in hospital settings, these tools are being used in hospitals much faster than it takes to establish formal regulations and guidelines for their use. Evaluations of the accuracy, completeness, and reproducibility of such advice across languages are still lacking. Therefore, the study aims to evaluate and compare the accuracy, completeness, safety, and reproducibility of three widely used AI chatbots—ChatGPT, Google Gemini, and Microsoft Copilot—when providing antidiabetic medication adjustment advice during Ramadan in both English and Arabic. Methods Twenty-three standardized clinical scenarios covering common antidiabetic regimens were presented to each chatbot in both English and Arabic. Each query was repeated to evaluate reproducibility, resulting in 276 responses scored. Responses were assessed against the International Diabetes Federation–Diabetes and Ramadan (IDF-DAR) Guidelines using a 0–2 accuracy scale, a 0–4 completeness scale, and a 0–3 safety scale. Results Overall, 77% of responses were fully consistent with the guideline, 12% were partially consistent, and 11% (30/276) contained clinically harmful or contradictory advice; harmful responses were about twice as common in Arabic as in English (14% vs. 8%). Completeness and safety were high, with medians at the observed ceiling. In the generalized linear mixed models, chatbots did not differ significantly in accuracy, completeness, or safety, and there was no significant main effect of language or chatbot × language interaction; the strongest signals were a chatbot effect on completeness (p = 0.068) and a language effect on safety (p = 0.064), both non-significant. Two-week reproducibility was fair for accuracy (weighted κ = 0.20, p = 0.009) and completeness (κ = 0.29, p = 0.001) and showed a very low κ in the safety scale (κ = 0.02, p = 0.81). Conclusions AI chatbots demonstrated comparable performance in delivering guideline-based advice for diabetes management during Ramadan, with no significant differences in accuracy, completeness, or safety. While most responses aligned with the IDF-DAR guideline, some harmful recommendations persisted, and response consistency fluctuated over time. These results suggest that AI chatbots should serve as a supplementary resource rather than a substitute for professional medical advice.
Background: Artificial intelligence chatbots are increasingly used to obtain medical and drug-related information, but their accuracy for clinical use remains uncertain. Objective: To evaluate and compare the performance of three large language models—ChatGPT, Gemini, and Grok—in responding to standardised drug-related...
S. Dhohan, Gagan D. Urs, K. Sneha et al.· Advances in Research· 0 citations
The study results demonstrated the chatbots' responses were typically suboptimal; albeit, a significant minority prompted a follow up to review the chatbot response, and the study conforms with the Declaration of Helsinki.
Duncan Yorkston, Tracey Borrie, Paul K. L. Chin· Journal of Pharmacy Practice...· 0 citations
Multilingual AI chatbots demonstrate a boost in healthcare efficiency, a reduction in language barriers, and the promotion of health equity, but exhibit challenges regarding validation, workflow integration, and evaluation standards, along with ethical issues such as privacy and bias.
R. Sharanesha, Deepti Virupakshappa, A. Abushanan et al.· Informatics· 0 citations
This study provides the first prototype of an AI-driven chatbot specifically designed for MAT professionals, demonstrating feasibility of integrating advanced AI technologies to address information access barriers in addiction treatment.
Sandra C. Nwobi, Zainab Loukil, Abbas Jawahar· Frontiers in Digital Health· 0 citations
Background/Objectives: Rosacea is a chronic inflammatory skin disease that requires long-term management and continuous patient education regarding triggers, skincare practices, and treatment adherence. In recent years, patients have increasingly turned to online platforms and artificial intelligence (AI)-based chatbot...
Mahmut Talha Uçar, Ecem Bostan, Tulay Ortabag et al.· Diagnostics· 0 citations
Background Patients with painful diabetic peripheral neuropathy (PDPN) increasingly use generative artificial intelligence chatbots for information on symptoms, treatment, foot care, and when to seek professional help. Their usefulness depends on safety, accuracy, guideline concordance, actionability, and readability....
Li-Na Gu, Zhao-Le Gong, Xiao-Li Qian et al.· Frontiers in Public Health· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.