Aug 2026· Annals of the Royal College of Surgeons of England· 0 citations· 36 references
Medicine
TL;DR
There may be significant differences in how effectively LLMs support patients with surgical queries, particularly in areas needing detailed explanation, and usually required minimal clarification in areas needing detailed explanation.
Abstract
INTRODUCTION
Large language models (LLMs) are increasingly used as sources of information across many fields, including healthcare. As patients turn to these models for health-related queries, evaluating the accuracy and reliability of their responses is essential. Peri-acetabular osteotomy (PAO) is performed on younger patients - a group more likely to use digital tools like LLMs for health information. This study assesses the accuracy and readability performance of two leading LLMs, ChatGPT and Google Gemini, in addressing common patient questions on PAO.
Methods
A panel of fellowship-trained PAO surgeons created ten commonly asked patient questions based on real-world experience. Responses from each LLM were assessed by the same three surgeons, blinded to response origin, using a 5-point Likert scale to evaluate clarity, accuracy, and completeness. Readability was measured with Flesch-Kincaid Reading scores.
Results
ChatGPT outperformed Gemini with an average score of 4.17 vs 3.13 (t = -3.08, p = 0.006). ChatGPT's responses were often rated higher for completeness and clarity, particularly in areas needing detailed explanation, and usually required minimal clarification. Gemini sometimes lacked specificity or included minor inaccuracies that reduced its perceived reliability. Both LLMs produced responses with similar "difficult" Flesch Reading Ease scores.
Conclusions
There may be significant differences in how effectively LLMs support patients with surgical queries. ChatGPT more consistently met expert standards for clarity and thoroughness. As LLM usage expands, ChatGPT may aid patient education on hip surgery, supporting consultations, informed decisions and postoperative guidance.
Background: Patellofemoral pain syndrome (PFPS) is a common cause of anterior knee pain, and patients increasingly use large language models (LLMs) to obtain general medical information. However, the quality, reliability, and readability of LLM-generated responses to patient-oriented questions regarding PFPS remain unc...
Oktay Polat, Berk Koncalıoğlu, Mert Gündoğdu et al.· Healthcare· 0 citations
ChatGPT-4o and ChatGPT-5 provide generally satisfactory yet non-comprehensive, limited-quality information at a level above tenth-grade regarding hallux rigidus fusion surgery.
Kamil Balaban, Mehmet Batu Ertan, Mahmut Kalem· Digital Health· 0 citations
Patients with hallux valgus increasingly seek health information through consumer-facing artificial intelligence (AI)–driven patient education tools, particularly large language model–based conversational agents. Although these tools offer rapid and accessible responses, concerns remain regarding the reliability,...
A. Koluman, Ebru Aloğlu Çiftçi, Mehmet Utku Çiftçi et al.· BMC Medical Informatics and...· 0 citations
AI models demonstrate generally high accuracy in medical examinations; however, they struggle with interpreting clinical context, recognising atypical medical scenarios and answering questions involving visual content.
Bünyamin Arı· Health Informatics Journal· 0 citations
LLMs can support communication and education following spine surgery when used with structured prompting when used with structured prompting and ChatGPT and Claude showed the highest correctness and completeness, particularly for practitioner-directed answers.
S. Wegmann, T. Rosenkranz, Philipp Egenolf et al.· European spine journal· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.