Jul 2026· Archives of Orthopaedic and Trauma Surgery· Vol 146· 0 citations· 35 references
Medicine
Abstract
This descriptive study aimed to longitudinally evaluate the performance of contemporary large language models - ChatGPT-5, Gemini 2.5 Flash, and Grok-3 - on orthopaedic clinical multiple-choice tasks, benchmarked against pooled clinician consensus. A secondary aim was to assess whether recent advances in generative AI translated into improved alignment with clinician consensus compared with previous AI models. A total of 97 multiple-choice clinical cases spanning eight orthopaedic subspecialties were sourced from OrthoBullets and previously benchmarked against aggregated responses from thousands of practising clinicians. Using identical methodology to our 2023 study of ChatGPT-3.5, ChatGPT-4, and Bard, each model was prompted with standardised case stems and response options. The primary outcome was the proportion of AI responses matching the most popular clinician response; secondary analyses assessed agreement within 10% and 20% of clinician consensus, performance on ‘controversial’ (< 25% margin) questions, and inter-model concordance using Cohen’s kappa coefficients. Gemini 2.5 Flash achieved the highest alignment with clinician consensus (69.1%), followed by Grok-3 (66.0%) and ChatGPT-5 (58.8%). None of the LLMs refused to respond to any prompts, representing a reduction from 7.2% from our 2023 study. Subspecialty analysis demonstrated that Gemini 2.5 Flash performed best in Hand and Paediatric domains, while Grok-3 excelled in Reconstruction, Trauma, and ‘controversial’ cases. Inter-model agreement was highest between Grok-3 and Gemini 2.5 Flash (κ = 0.678), indicating improved consistency compared with prior-generation systems. Contemporary LLMs can be promising adjuncts for orthopaedic education by simulating peer reasoning and offering structured explanations in non-critical settings. Despite incremental gains in reasoning capability compared to previous AI models, contemporary LLMs remain unsuitable for independent clinical use. Future research should develop hybrid clinician–AI workflows and longitudinal benchmarks to distinguish true reasoning improvements from memorisation.
OE demonstrated superior CSCG concordance and substantially lower susceptibility to hallucinations than ChatGPT, although accuracy declined when outputs were reformatted through ChatGPT, suggesting cross-model contamination.
F. Avrumova, Michael D. D. Cesar, Giuseppe Loggia et al.· The spine journal· 0 citations
Gemini-2.5-Flash provided the most reliable responses to common patient questions about rTKA generated by leading LLMs, highlighting the need for supervised integration of LLMs in patient education.
Mehmet Utku Çiftçi, A. Koluman, Ebru Aloğlu Çiftçi et al.· Knee (Oxford)· 1 citation
ChatGPT 5.2 is a valuable AI–assisted chatbot that facilitates patient education by providing responses regarding sleeve gastrectomy that are generally accurate and acceptable, but the categorization of 4–16% of the responses as "Incorrect," the overall difficult readability levels, and the significant variability obse...
Furkan Türkoğlu, Elif Nur Gencer, Emre Erdoğan· Archives of Current Medical...· 0 citations
Patients with hallux valgus increasingly seek health information through consumer-facing artificial intelligence (AI)–driven patient education tools, particularly large language model–based conversational agents. Although these tools offer rapid and accessible responses, concerns remain regarding the reliability,...
A. Koluman, Ebru Aloğlu Çiftçi, Mehmet Utku Çiftçi et al.· BMC Medical Informatics and...· 0 citations
Background: Artificial intelligence (AI) has expanded rapidly across orthopaedic practice, yet routine clinical adoption remains limited despite strong technical performance. This narrative review examines why a persistent gap separates technical maturity from clinical maturity across the orthopaedic patient care pathw...
Rafael De Nigris González, P. Mello· Journal of Clinical Medicine· 0 citations
Background: Generative artificial intelligence (AI), including large language models (LLMs), has been increasingly explored in orthopedic surgery; however, its application within total hip and knee arthroplasty (THA/TKA) has not been clearly characterized. Therefore, we performed a systematic review to further evaluate...
Ivan A. Garces, Andres G. Wong, J. Brutti et al.· JB & JS open access· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.