Aug 2026· Journal of Orthopaedic Surgery and Research· 0 citations
TL;DR
AI-generated PILs offer brevity but do not consistently improve readability, with some indices suggesting increased complexity, and clinician oversight and further validation are essential to ensure AI-generated materials enhance, rather than hinder, patient understanding and engagement.
Abstract
Effective patient education is critical in orthopaedic care, influencing satisfaction, adherence, and outcomes. Artificial intelligence (AI), particularly large language models (LLMs), could offer the potential to improve patient information leaflets (PILs), but this remains underexplored. This study aimed to evaluate the readability of AI-generated orthopaedic PILs compared to UK professional orthopaedic society materials using objective metrics.
A retrospective quantitative study was conducted comparing PILs from nine UK orthopaedic subspecialty societies with matched AI-generated counterparts created using ChatGPT 4.5. AI responses were generated using simple, single lined patient-style prompts to simulate real-world queries. PILs were categorised as either condition-based, procedure-based, or general information leaflets. Readability was assessed using validated metrics including Flesch-Kincaid Grade Level (FKGL) and Reading Age, FORCAST, New Dale-Chall, SMOG, Gunning Fog Index, and Flesch Reading Ease (FRE). Word counts were also analysed. Grade levels were interpreted according to U.S. educational standards. Statistical comparisons between AI and human-generated materials were performed using appropriate parametric and non-parametric tests, with statistical significance set at
p
< 0.05.
Across 134 orthopaedic PILs, AI-generated materials were consistently shorter in word count across all categories (
p
< 0.01). Despite no significant differences in FKGL for General and Procedure based PILs, AI-generated Condition PILs demonstrated significantly higher FKGL (10 [9.2–10.6] vs. 8.7 [7.8–9.6],
p
< 0.01). FRE was consistently lower in AI-generated texts across all categories (
p
< 0.01), suggesting reduced accessibility. AI materials also demonstrated significantly higher FORCAST and New Dale-Chall Grade Levels across all categories (all
p
< 0.01), indicating greater reading complexity.
AI-generated PILs offer brevity but do not consistently improve readability, with some indices suggesting increased complexity. While AI holds promise, clinician oversight and further validation are essential to ensure AI-generated materials enhance, rather than hinder, patient understanding and engagement.
Patients with hallux valgus increasingly seek health information through consumer-facing artificial intelligence (AI)–driven patient education tools, particularly large language model–based conversational agents. Although these tools offer rapid and accessible responses, concerns remain regarding the reliability,...
A. Koluman, Ebru Aloğlu Çiftçi, Mehmet Utku Çiftçi et al.· BMC Medical Informatics and...· 0 citations
There may be significant differences in how effectively LLMs support patients with surgical queries, particularly in areas needing detailed explanation, and usually required minimal clarification in areas needing detailed explanation.
TP Davis, B. Guevel, K. Logishetty et al.· Annals of the Royal College...· 0 citations
AI chatbots can generate accurate and comprehensive responses to common tracheostomy care questions, demonstrating potential to support patient education, but they continue to lack guaranteed, verifiable sourcing.
Keer Zhang, Lauran K. Evans, Desiree Delavary et al.· Otolaryngology Head & Neck S...· 0 citations
ChatGPT-4o and ChatGPT-5 provide generally satisfactory yet non-comprehensive, limited-quality information at a level above tenth-grade regarding hallux rigidus fusion surgery.
Kamil Balaban, Mehmet Batu Ertan, Mahmut Kalem· Digital Health· 0 citations
Abstract Background Patient visit summaries (PVS) are patient-facing documents intended to reinforce communication and promote patient education after clinical encounters. Despite national recommendations that patient education materials be written at or below a sixth-grade reading level, most orthopedic materials subs...
Edward Lee Major, Vivek P. Shah, Amber N. Carroll et al.· Journal of Medical Internet...· 0 citations
Patient comprehension of surgical information is often limited by medical jargon, health literacy barriers, and time constraints. Large language models (LLMs) such as ChatGPT‐5, Claude‐4, and Google AI search offer interactive context specific dialogue that may aid to overcome these limitations. To date, no study has...
Darcy Noll, T. Milton, P. Stapleton et al.· Trends in Urology & Men'...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.