Quality and readability of AI Chatbot responses to frequently asked questions from patients undergoing progressive collapsing foot deformity surgery: a comparative study of ChatGPT, Perplexity, and Gemini.
Aug 2026· Orthopaedics & Traumatology: Surgery & Research· pp.
104814
· 0 citations· 30 references
Medicine
TL;DR
AI chatbots produce generally accurate baseline information on PCFD surgery, with Perplexity showing significantly higher expert-rated accuracy and clarity than the other platforms-contrary to the hypothesis of comparable accuracy-while readability remains uniformly inadequate for all platforms, as hypothesized.
Abstract
Background
Patients increasingly turn to AI chatbots for medical information, including before complex orthopaedic procedures such as progressive collapsing foot deformity (PCFD) surgery. Whether these tools deliver content of sufficient quality and accessibility for preoperative patient education remains unclear, particularly across competing platforms. This study addressed three questions: (1) Do ChatGPT, Perplexity AI and Google Gemini differ in the accuracy, comprehensiveness and clarity of their responses to PCFD-related patient questions? (2) Do these platforms produce content meeting recommended readability thresholds for patient education? (3) Does the level of agreement among blinded foot and ankle surgeons rating the quality of chatbot responses vary depending on the platform used?
HYPOTHESIS
The three AI chatbot platforms produce responses of comparable accuracy but differ significantly in readability, with none reaching the recommended readability thresholds for patient education materials.
PATIENTS AND
Methods
Cross-sectional comparative study. Twenty frequently asked questions regarding PCFD, covering disease understanding, conservative management, surgical planning and postoperative recovery, were submitted verbatim to ChatGPT (GPT-4o mini), Perplexity AI and Google Gemini (free versions, March 25, 2026). The 60 resulting responses were rated by three blinded foot and ankle surgeons on three 5-point Likert scales (accuracy, comprehensiveness, clarity). Readability was assessed using the Flesch Reading Ease (FRE) and Flesch-Kincaid Grade Level (FKGL). Inter-rater agreement used Kendall's W; differences between platforms were analysed using Kruskal-Wallis tests with Bonferroni-corrected pairwise comparisons.
Results
All platforms produced responses rated accurate to very accurate. Perplexity achieved significantly higher accuracy than ChatGPT (p = 0.0003) and higher accuracy and clarity than Gemini (p = 0.0065 and p = 0.0075). No platform reached the recommended FRE ≥ 60 or FKGL ≤ 6 thresholds: median FKGL ranged from 13.1 (ChatGPT) to 21.4 (Perplexity), with ChatGPT producing the most readable and Perplexity the least readable content (p < 0.001). Inter-rater agreement was fair to substantial across platforms, lowest for Perplexity.
Discussion
AI chatbots produce generally accurate baseline information on PCFD surgery, with Perplexity showing significantly higher expert-rated accuracy and clarity than the other platforms-contrary to our hypothesis of comparable accuracy-while readability remains uniformly inadequate for all platforms, as hypothesized. These tools may serve as a supplementary source of information, but their inadequate readability suggests they are not yet suited to replace tailored, surgeon-led patient education.
LEVEL OF EVIDENCE
III; cross-sectional comparative study.
ChatGPT-4o and ChatGPT-5 provide generally satisfactory yet non-comprehensive, limited-quality information at a level above tenth-grade regarding hallux rigidus fusion surgery.
Kamil Balaban, Mehmet Batu Ertan, Mahmut Kalem· Digital Health· 0 citations
Patients with hallux valgus increasingly seek health information through consumer-facing artificial intelligence (AI)–driven patient education tools, particularly large language model–based conversational agents. Although these tools offer rapid and accessible responses, concerns remain regarding the reliability,...
A. Koluman, Ebru Aloğlu Çiftçi, Mehmet Utku Çiftçi et al.· BMC Medical Informatics and...· 0 citations
BACKGROUND
Patellofemoral Pain Syndrome (PFPS) is a highly prevalent musculoskeletal condition affecting young adults and athletes. Patients increasingly turn to AI chatbots for medical information, yet the reliability, safety, and readability of these tools for PFPS remain unclear.
OBJECTIVE
To evaluate accuracy, cl...
LLMs can support communication and education following spine surgery when used with structured prompting when used with structured prompting and ChatGPT and Claude showed the highest correctness and completeness, particularly for practitioner-directed answers.
S. Wegmann, T. Rosenkranz, Philipp Egenolf et al.· European spine journal· 0 citations
OBJECTIVE
To compare the quality and readability of responses from five generative artificial intelligence chatbot platforms to clinician-oriented questions on open temporomandibular joint (TMJ) surgery against guideline-based reference answers.
MATERIAL AND METHODS
Forty questions across eight domains were submitted...
While ChatGPT provided generally acceptable clinical information, variability in readability and expert ratings raises patient safety concerns and AI-generated health content should be reviewed and tailored to older adults' health literacy needs before clinical use.
Merve Arı, N. Ilçin, Hatice Yağcıoğlu et al.· Revista da Associação Médica...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.