Skip to content
Open access

Mapping Gaps and Improvement Targets in Large Language Model-Generated Melanoma Patient Education in a Non-English Setting

Aug 2026 · European Journal of Therapeutics · 0 citations · 26 references

TL;DR

How well large language models (LLM) handle Turkish melanoma patient education varies widely from one model to the next, and findings suggest that LLM-generated Turkish melanoma materials may be useful as preliminary educational drafts.

Abstract

Objective: Large language models (LLMs) are increasingly being used to develop medical education materials; however, it remains unclear how reliable, readable, or guideline-compliant the content generated by these models is for non-English-speaking patient groups. We evaluated the quality of Turkish melanoma patient education texts generated by seven frontier LLMs. Methods: A standardized 22-item Turkish prompt, built from international melanoma guidelines, was put to seven models in zero-shot sessions: ChatGPT 4.0 Turbo, Gemini 2.0 Flash, Claude 3.7 Sonnet, Grok 3, Qwen 2.5 Plus, DeepSeek R1, and Mistral Large 2. Each output was rated for readability (Ateşman Index), understandability, how clearly medical terminology was explained, scientific reliability (DISCERN instrument), empathy, and adherence to a 31-item guideline-based checklist. Model comparisons were summarized descriptively, using model-level absolute scores, score ranges, and rankings. Results: Model performance differed across readability, understandability, reliability, empathy, and guideline-adherence domains. DeepSeek R1 led on both readability (81.6) and understandability (23.5/25). Guideline adherence was strongest for Grok 3 and DeepSeek R1, at 96.8% and 93.5%, respectively, and Grok 3, DeepSeek R1, and Gemini 2.0 Flash each scored above 90% on the normalized total DISCERN measure. DeepSeek R1 also recorded the highest empathy score (90%). Gemini 2.0 Flash had the lowest readability score and produced the longest output (Ateşman 65.8). None of the models provided citations or verifiable sources, so every model received the lowest possible DISCERN Source Reliability score; Mistral Large 2 showed the weakest overall performance. Conclusion: How well large language models (LLM) handle Turkish melanoma patient education varies widely from one model to the next. A few produced text that was clear, empathetic, and reasonably guideline-concordant, but the lack of verifiable citations and uneven guideline coverage remain genuine limitations. These findings suggest that LLM-generated Turkish melanoma materials may be useful as preliminary educational drafts.

Read PDF

Similar papers

Open access Sep 2026

Performance evaluation of large language models in bladder cancer patient education Q&A: a cross-sectional study

Background Bladder cancer ranks among the most prevalent urological tumors worldwide, with its global incidence continuing to rise steadily. Although patient education materials (PEMs) play a crucial role in enhancing disease comprehension and supporting joint clinical decision-making, current online resources frequent...

Dian Wan, You-Wen Li, Zheng Dong et al. · 0 citations
Aug 2026

Performance of Large Language Models in Oral Cancer Patient Education: An Evaluation of Reliability, Readability, and Patient Communication Quality

Evaluating the reliability and readability of the responses generated by four mainstream LLMs to questions related to oral cancer found no model showed consistently high performance across all dimensions or met recommended readability standards.

Bo Zhang, Weidi Shi, Ying Zhang · 0 citations
Sep 2026

Large Language Models for Breast Cancer Education: A Comparative Analysis of Quality, Reliability and Readability.

BackgroundPatients increasingly consult artificial intelligence (AI) tools for breast cancer information. While Large Language Models (LLMs) enhance information accessibility, their accuracy, reliability, and alignment with patient health literacy remain critical concerns. This study compared the quality, reliability,...

Burak Altunpak · 0 citations
Review Open access Sep 2026

Benchmarking Large Language Model Performance in Generating and Assessing Radiology Objective Structured Clinical Examination

High-quality radiology assessment questions are essential for education competency evaluation but labor-intensive to create. To compare four large language models (LLMs) in generating and evaluating radiology objective structured clinical examination (OSCE)–style questions and responses. Fifty Radio...

Ankush Ankush, Samriddhi Burman, Sydney Smith et al. · 0 citations
#large language models Review Sep 2026

A Real-World Evaluation of Large Language Model-Generated Hospital Courses in Pediatrics.

BACKGROUND Large language model (LLM)-generated hospital courses are increasingly integrated into electronic health records (EHRs), yet their accuracy and safety in pediatric populations remain poorly characterized. OBJECTIVE To evaluate the accuracy, text quality, and perceived potential harm of EHR-integrated and L...

Jasmine E. Kim, J. Hron, Daniel J Kats et al. · 0 citations
Review Open access Aug 2026

A locally deployed large language model for pathology-informed and nurse-reviewed communication support in bladder cancer immunotherapy

Background Immunotherapy plays an important role in bladder cancer care, requiring ongoing patient education, symptom monitoring, and communication of pathology- and biomarker-related information. Locally deployed large language models (LLMs) may support these nurse-led activities, but their safety and clinical usabili...

Suqing Diao, Jun Zhao, Xu-Zhong Liu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.