Four general-purpose large language models (ChatGPT-5, Claude 4, Grok 4 and Gemini 2.5) show comparable performance in specialised total knee arthroplasty clinical questions.
PURPOSE To evaluate and compare the performance of four general-purpose large language models (LLMs) (ChatGPT-5, Claude 4, Grok 4 and Gemini 2.5) in answering specialised clinical questions related to total knee arthroplasty (TKA) derived from the World Expert Meeting in Arthroplasty (WEMA). METHODS This is a cross-s...