Performance of five large language models in frozen shoulder patient education: a multidimensional evaluation of readability, quality, clinical alignment, and safety
Abstract
This study compared the educational suitability, overall quality, translation-based readability, clinical intent alignment, and safety of frozen shoulder patient education materials generated by GPT-5, DeepSeek, Doubao, Wenxin Yiyan, and Tongyi Qianwen. Twenty clinician-curated questions covering five content categories were submitted once to each model, producing 100 first-pass responses. C-PEMAT and GQS assessed educational quality, seven English readability formulas were applied to consensus translations of the Chinese outputs, and two preliminary clinician-developed measures assessed clinical key-point coverage and safety-related caution. C-PEMAT and GQS differed significantly among models (both P < 0.001), with GPT-5 obtaining the highest scores, followed by DeepSeek and Doubao. Quality scores did not differ significantly across content categories, whereas all readability indices did. C-PEMAT and GQS were moderately correlated (r = 0.68), but quality measures were weakly associated with most readability indices. GPT-5 and DeepSeek showed relatively favorable Clinical Intent Alignment scores, while GPT-5 had the highest descriptive Clinical Safety Score. These findings are specific to the tested access date, platforms, prompts, and model versions. LLM-generated materials may serve as auxiliary drafts, but clinician review remains necessary for medical accuracy, safety, readability, and patient-specific appropriateness.