Skip to content
Open access

Safety and quality of public chatbots for lung cancer prognostic information: a comparative evaluation

Aug 2026 · Frontiers in Public Health · Vol 14 · 0 citations · 30 references
Medicine

TL;DR

Public-facing chatbots may support general patient education but should not replace individualized clinician-led prognostic communication as public-facing chatbots differed substantially in safety, reliability, communication quality, and readability.

Abstract

Objective To compare the safety, accuracy, empathy, reliability, information quality, and readability of five publicly accessible large language model chatbots when answering patient-facing lung cancer prognostic questions under standardized single-turn English prompting. Methods In this Chatbot Health Advice Reporting Transparency-guided cross-sectional evaluation, 53 standardized English prompts were submitted once to ChatGPT, Gemini, Copilot, DeepSeek, and Doubao through official web interfaces during April 1–21, 2026. Five blinded raters assessed 265 responses for safety, accuracy, empathy, DISCERN, EQIP, JAMA benchmark criteria, Global Quality Scale, and readability. Paired repeated-measures analyses were used. Results Inter-rater agreement was good to excellent. Safety differed significantly across models (Cochran’s Q = 14.089, df = 4, p = 0.007). Gemini generated the highest proportion of safe responses (48/53, 90.6%), whereas DeepSeek generated the lowest (33/53, 62.3%). The only adjusted pairwise safety difference that remained significant was Gemini versus DeepSeek (adjusted p = 0.023). Accuracy, empathy, reliability, information quality, and readability also differed significantly across models (all p < 0.001). Gemini showed the most favorable descriptive profile for safety, accuracy, empathy, and reliability, while Copilot produced the most readable responses. Conclusion Public-facing chatbots differed substantially in safety, reliability, communication quality, and readability. These findings are time-, interface-, and prompt-dependent. Chatbots may support general patient education but should not replace individualized clinician-led prognostic communication.

Read PDF

Similar papers

#small language model Open access Sep 2026

Safety, accuracy, empathy, information quality, and readability of publicly accessible LLM-based chatbots for traumatic brain injury and concussion questions: a cross-sectional comparative study

Background Large language model (LLM)-based chatbots are increasingly used by the public to obtain health information, but their performance in answering questions related to traumatic brain injury (TBI) and concussion remains unclear. This study evaluated five publicly accessible LLM-based chatbots across safety, accu...

Xin Zuo, Huan Zuo, Min Zhang et al. · 0 citations
#large language models Review Open access Sep 2026

A cross-sectional evaluation of large language model chatbot interfaces for patient-facing herpes zoster information: safety, information quality, and readability

Background Large language model (LLM) chatbots are increasingly used to obtain health information. However, fluent and clinically plausible responses may still contain safety-relevant omissions, inadequate source attribution and disclosure, or difficult-to-read text. Objective To evaluate the safety, accuracy, empathy,...

Da-Lian Liang, Cong Mai, Zheng-Kun Zhang et al. · 0 citations
Open access Aug 2026

Safety, accuracy, empathic communication, information quality, and readability of five large language model interfaces answering public questions about interstitial cystitis/bladder pain syndrome

Publicly accessible LLM interfaces showed domain-specific differences when answering IC/BPS-related public questions and may support general IC/BPS education and question preparation but should not replace clinician-led evaluation or individualized medical advice.

Jiang-Tao Zhu, Zhen-Hua Zhao, Song Li et al. · 0 citations
Open access Sep 2026

Comparative evaluation of ChatGPT and gemini responses to patient-oriented questions on breast cancer

ChatGPT’s higher scores and shorter, more focused responses indicate that it may be a more efficient tool for addressing breast cancer–related patient questions, and Gemini’s acceptable overall performance demonstrates acceptable overall performance.

Yeliz Yılmaz Bozok, N. Acar, M. Atahan et al. · 0 citations
Open access Aug 2026

Safety, accuracy, empathy, reliability, and readability of large language model chatbot responses to public-facing vegetarian and vegan nutrition advice questions: a cross-sectional comparative study

Background Publicly accessible large language model (LLM) chatbots are increasingly used to seek nutrition advice. Although vegetarian and vegan nutrition advice is often perceived as low risk, it may involve supplementation, vulnerable life stages, chronic disease, and symptoms that require clinical assessment. Object...

Jiyong Zhou, Qiongjiao Zhou, Lijun Zhou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.