Skip to content
Open access

Artificial intelligence advancements for orthopaedic clinical reasoning: longitudinal assessment of newer models (ChatGPT-5, Grok-3, Gemini 2.5 Flash) compared to clinicians

Jul 2026 · Archives of Orthopaedic and Trauma Surgery · Vol 146 · 0 citations · 35 references
Medicine

Abstract

This descriptive study aimed to longitudinally evaluate the performance of contemporary large language models - ChatGPT-5, Gemini 2.5 Flash, and Grok-3 - on orthopaedic clinical multiple-choice tasks, benchmarked against pooled clinician consensus. A secondary aim was to assess whether recent advances in generative AI translated into improved alignment with clinician consensus compared with previous AI models. A total of 97 multiple-choice clinical cases spanning eight orthopaedic subspecialties were sourced from OrthoBullets and previously benchmarked against aggregated responses from thousands of practising clinicians. Using identical methodology to our 2023 study of ChatGPT-3.5, ChatGPT-4, and Bard, each model was prompted with standardised case stems and response options. The primary outcome was the proportion of AI responses matching the most popular clinician response; secondary analyses assessed agreement within 10% and 20% of clinician consensus, performance on ‘controversial’ (< 25% margin) questions, and inter-model concordance using Cohen’s kappa coefficients. Gemini 2.5 Flash achieved the highest alignment with clinician consensus (69.1%), followed by Grok-3 (66.0%) and ChatGPT-5 (58.8%). None of the LLMs refused to respond to any prompts, representing a reduction from 7.2% from our 2023 study. Subspecialty analysis demonstrated that Gemini 2.5 Flash performed best in Hand and Paediatric domains, while Grok-3 excelled in Reconstruction, Trauma, and ‘controversial’ cases. Inter-model agreement was highest between Grok-3 and Gemini 2.5 Flash (κ = 0.678), indicating improved consistency compared with prior-generation systems. Contemporary LLMs can be promising adjuncts for orthopaedic education by simulating peer reasoning and offering structured explanations in non-critical settings. Despite incremental gains in reasoning capability compared to previous AI models, contemporary LLMs remain unsuitable for independent clinical use. Future research should develop hybrid clinician–AI workflows and longitudinal benchmarks to distinguish true reasoning improvements from memorisation.

Read PDF

Similar papers

Review Aug 2026

Evolution of Generative Artificial Intelligence in Clinical Practice: Comparative Performance of OpenEvidence 2.0 and ChatGPT-4o.

OE demonstrated superior CSCG concordance and substantially lower susceptibility to hallucinations than ChatGPT, although accuracy declined when outputs were reformatted through ChatGPT, suggesting cross-model contamination.

F. Avrumova, Michael D. D. Cesar, Giuseppe Loggia et al. · 0 citations
Open access Sep 2026

Artificial Intelligence in Bariatric Patient Education: A Multi-rater Evaluation of Reliability, Readability, and Clinical Validity of ChatGPT 5.2

ChatGPT 5.2 is a valuable AI–assisted chatbot that facilitates patient education by providing responses regarding sleeve gastrectomy that are generally accurate and acceptable, but the categorization of 4–16% of the responses as "Incorrect," the overall difficult readability levels, and the significant variability obse...

Furkan Türkoğlu, Elif Nur Gencer, Emre Erdoğan · 0 citations
Review Open access Jul 2026

Evaluating the reliability, quality, and readability of AI-generated patient education on hallux valgus: a comparative study of large language models

Patients with hallux valgus increasingly seek health information through consumer-facing artificial intelligence (AI)–driven patient education tools, particularly large language model–based conversational agents. Although these tools offer rapid and accessible responses, concerns remain regarding the reliability,...

A. Koluman, Ebru Aloğlu Çiftçi, Mehmet Utku Çiftçi et al. · 0 citations
Review Open access Aug 2026

Artificial Intelligence in Orthopaedics: Current Evidence and Clinical Translation Across the Patient Care Pathway

Background: Artificial intelligence (AI) has expanded rapidly across orthopaedic practice, yet routine clinical adoption remains limited despite strong technical performance. This narrative review examines why a persistent gap separates technical maturity from clinical maturity across the orthopaedic patient care pathw...

Rafael De Nigris González, P. Mello · 0 citations
Review Open access Jul 2026

Generative Artificial Intelligence in Hip and Knee Arthroplasty: A Systematic Review of Emerging Clinical Applications in Patient Communication and Education, Documentation, and Decision Support

Background: Generative artificial intelligence (AI), including large language models (LLMs), has been increasingly explored in orthopedic surgery; however, its application within total hip and knee arthroplasty (THA/TKA) has not been clearly characterized. Therefore, we performed a systematic review to further evaluate...

Ivan A. Garces, Andres G. Wong, J. Brutti et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.