Skip to content

Performance of Large Language Models in Oral Cancer Patient Education: An Evaluation of Reliability, Readability, and Patient Communication Quality

Aug 2026 · Oral Health & Preventive Dentistry · Vol 24, pp. 659 - 667 · 0 citations · 29 references
Medicine

TL;DR

Evaluating the reliability and readability of the responses generated by four mainstream LLMs to questions related to oral cancer found no model showed consistently high performance across all dimensions or met recommended readability standards.

Abstract

Background As society is increasingly depending on large language models (LLMs) for health-related questions, it is essential to objectively evaluate the quality and accessibility of the oral cancer information they provide. Although LLMs occupy a growing space in digital health communication, it remains unknown whether the information they generate is both reliable and easy to read. Objective The purpose of this study was to evaluate the reliability and readability, respectively, of the responses generated by four mainstream LLMs (ChatGPT, Gemini, Perplexity, and DeepSeek) to questions related to oral cancer. Specifically, the present authors aimed to evaluate the reliability of responses to common oral cancer questions and assess whether the readability of responses meets established expectations. Methods Twenty-two commonly asked, patient-orientated oral cancer–related questions were developed through two predefined phases: Google Trends analysis and expert consultation with specialists in oral oncology. Each question was entered as an independent single-turn prompt into four LLMs: ChatGPT-5, Gemini 2.5, Perplexity Pro, and DeepSeek v3.2. The primary outcome was information reliability and quality, assessed using four standardized instruments: the DISCERN questionnaire, the Ensuring Quality Information for Patients (EQIP) tool, the Journal of the American Medical Association (JAMA) benchmark criteria, and the Global Quality Scale (GQS). The secondary outcome was readability, assessed using six established indices: the Automated Readability Index, Flesch Reading Ease Score, Gunning Fog Index, Flesch-Kincaid Grade Level, Coleman-Liau Index, and Simple Measure of Gobbledygook. Results Significant differences were observed among the four LLMs in DISCERN, EQIP, and JAMA scores (all P < 0.001), whereas no significant difference was found in GQS scores (P = 0.440). Perplexity Pro achieved the highest mean DISCERN score (46.36 ± 4.70), EQIP score (85.00 ± 0.00), GQS score (4.05 ± 0.58), and JAMA score (1.00 ± 0.00). However, all models produced responses above the recommended sixth-grade readability level. The mean FKGL scores ranged from 12.65 ± 3.07 for ChatGPT-5 to 15.65 ± 3.36 for Perplexity Pro, and the mean FRES scores ranged from 37.50 ± 14.27 for Perplexity Pro to 52.59 ± 12.47 for Gemini 2.5. Conclusion Current LLMs may support oral cancer patient education, but their use remains limited by variable information quality, insufficient transparency, and poor readability. Although Perplexity Pro performed better on several reliability-related metrics, no model showed consistently high performance across all dimensions or met recommended readability standards. Future LLM-based patient education tools should prioritise verifiable sourcing, guideline-based accuracy, risk communication, and plain-language adaptation.

View source

Similar papers

Sep 2026

Large Language Models for Breast Cancer Education: A Comparative Analysis of Quality, Reliability and Readability.

BackgroundPatients increasingly consult artificial intelligence (AI) tools for breast cancer information. While Large Language Models (LLMs) enhance information accessibility, their accuracy, reliability, and alignment with patient health literacy remain critical concerns. This study compared the quality, reliability,...

Burak Altunpak · 0 citations
#large language models Open access Sep 2026

A comparative study of large language models in responding to breast cancer–related questions

Breast cancer is one of the most common cancers among women worldwide. With advances in diagnostic and therapeutic technologies, more patients seek information about their condition. Large language models (LLMs) have attracted attention in healthcare for their ability to produce personalized health content. This stud...

Min-Xia Lin, Wen-Jun Liu, Zhi-Qiang Zeng · 0 citations
Open access Sep 2026

Performance evaluation of large language models in bladder cancer patient education Q&A: a cross-sectional study

Background Bladder cancer ranks among the most prevalent urological tumors worldwide, with its global incidence continuing to rise steadily. Although patient education materials (PEMs) play a crucial role in enhancing disease comprehension and supporting joint clinical decision-making, current online resources frequent...

Dian Wan, You-Wen Li, Zheng Dong et al. · 0 citations
Open access Aug 2026

Mapping Gaps and Improvement Targets in Large Language Model-Generated Melanoma Patient Education in a Non-English Setting

How well large language models (LLM) handle Turkish melanoma patient education varies widely from one model to the next, and findings suggest that LLM-generated Turkish melanoma materials may be useful as preliminary educational drafts.

Nıyazı Çetın, A. Atılan · 0 citations
Open access Sep 2026

Comparative evaluation of ChatGPT and gemini responses to patient-oriented questions on breast cancer

Background Breast cancer, the most common malignancy among women, remains a major global public health concern. With the rapid growth of artificial intelligence–based language models, it is essential to evaluate their potential roles in patient education. This study compared ChatGPT and Gemini in responding to patient-...

Yeliz Yılmaz Bozok, N. Acar, M. Atahan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.