Aug 2026· Journal of Visualized Experiments· Vol 234· 0 citations
Medicine
TL;DR
It is suggested that model selection influences expert-rated educational suitability and overall information quality, whereas topic complexity primarily affects readability.
Abstract
Large language models (LLMs) are increasingly used for patient health education, yet the readability and educational quality of LLM-generated information on trigeminal neuralgia (TN) have been insufficiently evaluated. This cross-sectional benchmarking study compared TN educational content generated by five publicly available LLMs (Doubao, DeepSeek, Wenxin Yiyan, Tongyi Qianwen, and GPT-5). Twenty frequently asked TN questions covering basic disease knowledge, etiology/risk factors, diagnosis, treatment, and prevention/rehabilitation were presented to each model using standardized prompts. Readability was assessed using seven established indices, educational suitability using the Patient Education Materials Assessment Tool for Understandability and Actionability (PEMAT), and overall information quality using the Global Quality Score (GQS). Two clinical experts independently evaluated all responses, with disagreements resolved by a senior adjudicator. Statistical analyses compared model performance, thematic differences, and correlations among the evaluation metrics. Significant differences were observed among the models for readability, PEMAT, and GQS scores. GPT-5 generated the most linguistically complex responses but achieved the highest ratings for educational suitability and information quality. In contrast, Wenxin Yiyan produced the most readable text but generally scored lower on PEMAT and GQS. Content category influenced readability, with prevention/rehabilitation and etiology/risk-factor topics being more difficult to read, whereas PEMAT and GQS remained relatively consistent across themes. Readability indices showed strong internal consistency and weak-to-moderate positive correlations with PEMAT and GQS, while PEMAT and GQS demonstrated a moderate positive correlation. These findings suggest that model selection influences expert-rated educational suitability and overall information quality, whereas topic complexity primarily affects readability. Because patient comprehension, satisfaction, trust, health outcomes, factual accuracy, and clinical safety were not evaluated, these results should be interpreted as an expert-rated benchmarking analysis rather than evidence of clinical readiness.
Objective To evaluate the performance of large language models (LLMs) and expert clinicians in optimizing Chinese patient education materials (PEMs) for temporomandibular disorders (TMD) across readability, accuracy, actionability, and cultural adaptability, and to determine whether a human-AI collaboration model can a...
Jinghong Han, Yankang Shi, Lijiao Dai· Frontiers in Public Health· 0 citations
DeepSeek demonstrates a significant advantage in information reliability, particularly excelling in postoperative and follow-up management content, and ChatGPT shows a slight edge in the readability of surgical planning sections.
Ling Tian, Long-Yue Tang, Ming-Tao Yang et al.· Frontiers in Public Health· 0 citations
Background: Patellofemoral pain syndrome (PFPS) is a common cause of anterior knee pain, and patients increasingly use large language models (LLMs) to obtain general medical information. However, the quality, reliability, and readability of LLM-generated responses to patient-oriented questions regarding PFPS remain unc...
Oktay Polat, Berk Koncalıoğlu, Mert Gündoğdu et al.· Healthcare· 0 citations
Large language models (LLMs) are increasingly used by the public to obtain health information, but their ability to provide reliable information for achalasia remains unclear. We therefore compared six LLMs in answering public questions about achalasia, an uncommon esophageal motility disorder, focusing on safety,...
Jun-Zheng Li, Pan Zhou, Man Yang et al.· Frontiers in Public Health· 0 citations
Patients with hallux valgus increasingly seek health information through consumer-facing artificial intelligence (AI)–driven patient education tools, particularly large language model–based conversational agents. Although these tools offer rapid and accessible responses, concerns remain regarding the reliability,...
A. Koluman, Ebru Aloğlu Çiftçi, Mehmet Utku Çiftçi et al.· BMC Medical Informatics and...· 0 citations
BackgroundPatients increasingly consult artificial intelligence (AI) tools for breast cancer information. While Large Language Models (LLMs) enhance information accessibility, their accuracy, reliability, and alignment with patient health literacy remain critical concerns. This study compared the quality, reliability,...
Burak Altunpak· Surgical Innovation· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.