Jul 2026
Performance of multimodal large language models versus clinicians for radiographic knee osteoarthritis grading: A multiobserver study.
Although ChatGPT-5.0 outperformed ChatGPT-4o, both models remained inferior to clinicians in detailed KL grading, feature-level interpretation, and reproducibility and are not suitable for standalone radiographic KOA assessment.
A. Akdoğan, Efe Kemal Akdoğan, Mehmet Fatih Tumer et al.
· Skeletal Radiology · 0 citations