Skip to content
Open access

Comparing the readability of AI-generated and society-authored patient information leaflets in orthopaedics

Aug 2026 · Journal of Orthopaedic Surgery and Research · 0 citations

TL;DR

AI-generated PILs offer brevity but do not consistently improve readability, with some indices suggesting increased complexity, and clinician oversight and further validation are essential to ensure AI-generated materials enhance, rather than hinder, patient understanding and engagement.

Abstract

Effective patient education is critical in orthopaedic care, influencing satisfaction, adherence, and outcomes. Artificial intelligence (AI), particularly large language models (LLMs), could offer the potential to improve patient information leaflets (PILs), but this remains underexplored. This study aimed to evaluate the readability of AI-generated orthopaedic PILs compared to UK professional orthopaedic society materials using objective metrics. A retrospective quantitative study was conducted comparing PILs from nine UK orthopaedic subspecialty societies with matched AI-generated counterparts created using ChatGPT 4.5. AI responses were generated using simple, single lined patient-style prompts to simulate real-world queries. PILs were categorised as either condition-based, procedure-based, or general information leaflets. Readability was assessed using validated metrics including Flesch-Kincaid Grade Level (FKGL) and Reading Age, FORCAST, New Dale-Chall, SMOG, Gunning Fog Index, and Flesch Reading Ease (FRE). Word counts were also analysed. Grade levels were interpreted according to U.S. educational standards. Statistical comparisons between AI and human-generated materials were performed using appropriate parametric and non-parametric tests, with statistical significance set at p  < 0.05. Across 134 orthopaedic PILs, AI-generated materials were consistently shorter in word count across all categories ( p  < 0.01). Despite no significant differences in FKGL for General and Procedure based PILs, AI-generated Condition PILs demonstrated significantly higher FKGL (10 [9.2–10.6] vs. 8.7 [7.8–9.6], p  < 0.01). FRE was consistently lower in AI-generated texts across all categories ( p  < 0.01), suggesting reduced accessibility. AI materials also demonstrated significantly higher FORCAST and New Dale-Chall Grade Levels across all categories (all p  < 0.01), indicating greater reading complexity. AI-generated PILs offer brevity but do not consistently improve readability, with some indices suggesting increased complexity. While AI holds promise, clinician oversight and further validation are essential to ensure AI-generated materials enhance, rather than hinder, patient understanding and engagement.

Read PDF

Similar papers

Review Open access Jul 2026

Evaluating the reliability, quality, and readability of AI-generated patient education on hallux valgus: a comparative study of large language models

Patients with hallux valgus increasingly seek health information through consumer-facing artificial intelligence (AI)–driven patient education tools, particularly large language model–based conversational agents. Although these tools offer rapid and accessible responses, concerns remain regarding the reliability,...

A. Koluman, Ebru Aloğlu Çiftçi, Mehmet Utku Çiftçi et al. · 0 citations
Open access Aug 2026

Evaluating large language models in patient education: a comparative analysis addressing frequently asked questions in peri-acetabular osteotomy.

There may be significant differences in how effectively LLMs support patients with surgical queries, particularly in areas needing detailed explanation, and usually required minimal clarification in areas needing detailed explanation.

TP Davis, B. Guevel, K. Logishetty et al. · 0 citations
Aug 2026

Quality of AI-Generated Patient Education for Pre- and Post-Operative Tracheostomy Care.

AI chatbots can generate accurate and comprehensive responses to common tracheostomy care questions, demonstrating potential to support patient education, but they continue to lack guaranteed, verifiable sourcing.

Keer Zhang, Lauran K. Evans, Desiree Delavary et al. · 0 citations
Review Open access Aug 2026

Readability of AI-Generated Patient Visit Summaries in Orthopedic Surgery: Retrospective Analysis

Abstract Background Patient visit summaries (PVS) are patient-facing documents intended to reinforce communication and promote patient education after clinical encounters. Despite national recommendations that patient education materials be written at or below a sixth-grade reading level, most orthopedic materials subs...

Edward Lee Major, Vivek P. Shah, Amber N. Carroll et al. · 0 citations
Open access Aug 2026

Evaluating the Role of AI Chatbots in Patient Education for Benign Scrotal Surgeries

Patient comprehension of surgical information is often limited by medical jargon, health literacy barriers, and time constraints. Large language models (LLMs) such as ChatGPT‐5, Claude‐4, and Google AI search offer interactive context specific dialogue that may aid to overcome these limitations. To date, no study has...

Darcy Noll, T. Milton, P. Stapleton et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.