Skip to content
Review Open access

Evaluating AI-generated patient education materials for endometrial cancer surgery: a comparative analysis of response quality, reliability, and readability between ChatGPT and DeepSeek models

Aug 2026 · Frontiers in Public Health · Vol 14 · 0 citations · 43 references
Medicine

TL;DR

DeepSeek demonstrates a significant advantage in information reliability, particularly excelling in postoperative and follow-up management content, and ChatGPT shows a slight edge in the readability of surgical planning sections.

Abstract

Purpose This study aimed to evaluate and compare the quality, reliability, and readability of patient education materials on endometrial cancer surgery generated by ChatGPT (GPT-5) and DeepSeek (R1). Materials and methods This cross-sectional study analyzed the responses generated by ChatGPT and DeepSeek to totally 41 questions covering four domains: surgical planning, preoperative evaluation, postoperative care, and long-term follow-up. Reliability was assessed through the DISCERN and EQIP instruments, quality was evaluated by the Global Quality Score (GQS), and readability was analyzed by the Flesch Reading Ease Score (FRES), Gunning Fog Index (GFI), and Flesch-Kincaid Grade Level (FKGL). Statistical comparisons were performed by using paired t-tests and Wilcoxon signed-rank tests. Results The two large language models (LLMs) generated education materials of comparable quality, as reflected in GQS scores (median: DeepSeek vs. ChatGPT 5.00 vs. 4.67, p = 0.077). DeepSeek demonstrated statistically significantly higher reliability scores on both DISCERN and EQIP instruments (both p < 0.001). Readability scores (FRES, GFI) were similar between groups, while DeepSeek exhibited a higher FKGL (10.28 vs. 8.84, p < 0.001), indicating the greater text complexity. Subgroup analysis showed that DeepSeek performed better in terms of reliability in the postoperative care and long-term follow-up domains, while ChatGPT exhibited better readability in the surgical planning domain. Conclusion Both DeepSeek and ChatGPT can generate patient education text drafts that are commendable in their structural coherence and linguistic clarity. DeepSeek demonstrates a significant advantage in information reliability, particularly excelling in postoperative and follow-up management content. ChatGPT shows a slight edge in the readability of surgical planning sections. However, the text readability of both models exceeds the general public's health literacy level. This indicates that large language models can only serve as auxiliary tools for generating patient education materials. Their outputs must undergo review by clinical experts and readability optimization to ensure both accuracy and comprehensibility of the information.

Read PDF

Similar papers

Sep 2026

Large Language Models for Breast Cancer Education: A Comparative Analysis of Quality, Reliability and Readability.

BackgroundPatients increasingly consult artificial intelligence (AI) tools for breast cancer information. While Large Language Models (LLMs) enhance information accessibility, their accuracy, reliability, and alignment with patient health literacy remain critical concerns. This study compared the quality, reliability, and readability of breast cancer-related responses generated by ChatGPT and Gemini.MethodsIn this cross-sectional study, conducted between March 20 and March 31, 2026, 40 questions spanning diagnosis, treatment, genetics, and follow-up were submitted to ChatGPT-5.3 and Gemini 3.0 Flash. Three independent surgeons evaluated the responses in a double-blinded manner using the modified DISCERN (mDISCERN) for reliability and Global Quality Score (GQS) for content quality assessment. Readability was assessed via Flesch Reading Ease (FRES), Flesch-Kincaid Grade Level (FKGL), Gunning Fog Index (GFI), and Simple Measure of Gobbledygook (SMOG) indices.ResultsGemini demonstrated statistically significant superiority over ChatGPT in both GQS (4.35 ± 0.30 vs 3.90 ± 0.29; P < .001) and mDISCERN (3.63 ± 0.58 vs 3.02 ± 0.48; P < .001) scores. In readability analysis, Gemini exhibited higher FRES (53.90 vs 45.13) and lower FKGL (9.31 vs 10.91) values, indicating enhanced patient accessibility (P < .05). For both models, the "Diagnosis" category yielded the highest readability, whereas "Treatment" scored the lowest. Inter-rater reliability for mDISCERN was moderate (ICC = 0.595).ConclusionsGemini significantly outperforms ChatGPT in response quality, reliability, and linguistic accessibility for breast cancer education. However, both models exceed the recommended sixth-grade reading level, indicating suboptimal optimization for general health literacy. While LLMs serve as promising auxiliary tools, expert supervision and cross-validation remain mandatory to ensure patient safety.

Burak Altunpak · 0 citations
Aug 2026

Quality of AI-Generated Patient Education for Pre- and Post-Operative Tracheostomy Care.

AI chatbots can generate accurate and comprehensive responses to common tracheostomy care questions, demonstrating potential to support patient education, but they continue to lack guaranteed, verifiable sourcing.

Keer Zhang, Lauran K. Evans, Desiree Delavary et al. · 0 citations
Open access Aug 2026

Mapping Gaps and Improvement Targets in Large Language Model-Generated Melanoma Patient Education in a Non-English Setting

How well large language models (LLM) handle Turkish melanoma patient education varies widely from one model to the next, and findings suggest that LLM-generated Turkish melanoma materials may be useful as preliminary educational drafts.

Nıyazı Çetın, A. Atılan · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.