Skip to content
Open access

Large language models as sources of patient information on robotic knee arthroplasty: a comparative evaluation.

Jul 2026 · Knee (Oxford) · Vol 62, pp. 104579 · 1 citation · 30 references
Medicine

TL;DR

Gemini-2.5-Flash provided the most reliable responses to common patient questions about rTKA generated by leading LLMs, highlighting the need for supervised integration of LLMs in patient education.

Abstract

Background

Robotic-assisted total knee arthroplasty (rTKA) is increasingly used because of its surgical precision. However, inconsistent outcomes and high costs often lead patients to seek additional information from artificial intelligence (AI) tools. Large language models (LLMs) such as ChatGPT-4o, Gemini-2.5-Flash, and DeepSeek-V3 are commonly used, but their reliability and readability in orthopaedics remain unclear.

Objectives

To compare the reliability, usefulness, quality, and readability of responses to common patient questions about rTKA generated by leading LLMs.

Methods

Three LLMs answered 20 frequently asked patient questions (n = 20) identified through Google Trends and expert validation. Three orthopaedic specialists (n = 3) evaluated reliability, usefulness, and overall quality using validated scales, while readability was assessed with standard indices.

Results

Inter-rater reliability was good to excellent (ICC = 0.728-0.879). Gemini-2.5-Flash achieved significantly higher reliability and usefulness scores than ChatGPT-4o and DeepSeek-V3 (all p < 0.05). ChatGPT-4o and DeepSeek-V3 produced more readable but less accurate content, revealing an inverse relationship between reliability and readability.

Conclusions

Gemini-2.5-Flash provided the most reliable responses, highlighting the need for supervised integration of LLMs in patient education.

Read PDF

Similar papers

Open access Sep 2026

Comparative quality, accuracy, and readability of large language model responses to patient questions about robotic-assisted total knee arthroplasty.

PURPOSE To compare the information quality, accuracy, and readability of patient-directed responses generated by large language models (LLMs), including ChatGPT-o3, ChatGPT-5.2, Gemini 3, and DeepSeek, regarding robotic-assisted total knee arthroplasty (RA-TKA). METHODS Thirty frequently asked patient questions were...

U. Kolaç, Mazlum Veysel Sili, Orhan Mete Karademir et al. · 0 citations
Open access Aug 2026

Evaluating large language models in patient education: a comparative analysis addressing frequently asked questions in peri-acetabular osteotomy.

There may be significant differences in how effectively LLMs support patients with surgical queries, particularly in areas needing detailed explanation, and usually required minimal clarification in areas needing detailed explanation.

T. Davis, B. Guevel, K. Logishetty et al. · 0 citations
Open access Sep 2026

Evaluating artificial intelligence-generated clinical guidance for patellar dislocation: accuracy, readability and qualitative appraisal of DeepSeek-R1 responses

Large language models (LLMs) are increasingly used to answer patient questions, but the readability and information quality of patellar dislocation guidance remain unclear. A series of clinical questions covering etiopathogenesis, mechanisms, clinical manifestations, diagnosis, treatment, complications, and...

Shuo Liu, Xi-Yang Sun, Yun-Fei Ma et al. · 0 citations
Open access Sep 2026

Evaluation of large language model responses to expert questions in anterior implant dentistry: quality, accuracy, and readability

AI chatbots can generate information with potential clinical relevance in anterior implant dentistry; however, variability in informational reliability persists and expert supervision remains essential before integrating such tools into clinical education.

Dalndushe Abdulai, Raghıb Suradı, Mehran Moghbel · 0 citations
Review Open access Jul 2026

Generative Artificial Intelligence in Hip and Knee Arthroplasty: A Systematic Review of Emerging Clinical Applications in Patient Communication and Education, Documentation, and Decision Support

Background: Generative artificial intelligence (AI), including large language models (LLMs), has been increasingly explored in orthopedic surgery; however, its application within total hip and knee arthroplasty (THA/TKA) has not been clearly characterized. Therefore, we performed a systematic review to further evaluate...

Ivan A. Garces, Andres G. Wong, J. Brutti et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.