Skip to content
Review Open access

Diagnostic Performance of Large Language Models for Orthopedic-Related Rare Diseases and Their Impact on Physicians’ Diagnostic Accuracy: 2-Stage Comparative Evaluation Study Based on the Chinese Rare Disease Catalog

Jul 2026 · Journal of Medical Internet Research · Vol 28, pp. e92931 · 0 citations · 50 references
Medicine

TL;DR

In the 2-stage LLM-assisted workflow, LLM assistance was associated with higher diagnostic correctness in both physician groups, although seniority-related differences in the magnitude of benefit require evaluation in larger studies.

Abstract

Background Orthopedic-related rare diseases are difficult to diagnose because of their low prevalence, heterogeneous phenotypes, and fragmented knowledge. Large language models (LLMs) can serve as dynamic knowledge-support tools, but their diagnostic performance and effect on physicians’ decision-making remain unclear. Objective This study aims to compare the diagnostic performance of advanced LLMs for orthopedic-related rare diseases and to evaluate the effect of a 2-stage LLM-assisted diagnostic workflow on physicians’ diagnostic accuracy and subjective acceptance. Methods We selected 40 orthopedic-related rare diseases from the Chinese Rare Disease Catalog. A total of 4 general-purpose LLMs each generated 1 primary diagnosis and 5 differential diagnoses per case. Diagnostic accuracy, defined as a correct primary diagnosis, was compared using the Cochran Q test and pairwise McNemar tests with Bonferroni correction. A representative LLM was integrated into a 2-stage workflow involving 27 intermediate and 15 senior orthopedic physicians. Physicians first diagnosed all cases independently and then rediagnosed the same cases after reviewing nonauthoritative LLM suggestions. Physician diagnostic data were primarily analyzed using mixed-effects logistic regression at the individual-diagnosis level. Case-level group accuracy was additionally assessed using >50% and ≥2/3 accurate-physician thresholds. After both rounds, physicians completed an 8-item Likert-scale questionnaire assessing subjective acceptance and workflow perceptions. Results Claude Sonnet 4.5, ChatGPT-5.0, and Gemini 2.5 Pro each achieved 90% (36/40) primary-diagnosis accuracy, whereas DeepSeek-V3.2 achieved 67.5% (27/40; Cochran Q P<.001). Before LLM assistance, mean physician-level accuracy was 42.22% for intermediate physicians and 58.67% for senior physicians; after assistance, it increased to 68.80% and 83.33%, respectively. In the primary mixed-effects logistic regression analysis, physician seniority group and LLM assistance stage were significantly associated with diagnostic correctness (both P<.001), whereas the group-by-stage interaction was not significant (P=.10). Secondary case-level analyses using the >50% threshold showed improvement from 40% (16/40) to 67.5% (27/40) for intermediate physicians and from 57.5% (23/40) to 82.5% (33/40) for senior physicians, with similar findings using the ≥2/3 threshold. Cases accurately diagnosed by all 3 agents increased from 16 to 27. The questionnaire showed high internal consistency (Cronbach α=0.902) and generally positive attitudes, with no significant differences between physician groups (P=.11 to P=.78). Conclusions LLMs achieved high diagnostic accuracy for orthopedic-related rare diseases. In the 2-stage LLM-assisted workflow, LLM assistance was associated with higher diagnostic correctness in both physician groups, although seniority-related differences in the magnitude of benefit require evaluation in larger studies. Senior physicians retained higher diagnostic correctness than intermediate physicians. Secondary case-level analyses suggested attenuation of group-level gaps in case-recognition patterns. Physicians reported broadly positive workflow perceptions. Given the same-day repeated-case design and potential short-term recall bias, these exploratory findings should be interpreted cautiously and warrant prospective randomized, crossover, washout-period, independent-case, or real-world evaluations of LLM-assisted diagnostic workflows in orthopedics.

Read PDF

Similar papers

Open access Aug 2026

Stepwise Diagnostic Evaluation of Chinese Large Language Models: Comparative Study of Common and Rare Diseases

Evaluated large language models for common diseases and rare diseases using clinical vignettes within a hypothetico-deductive framework demonstrated relatively strong diagnostic performance for common diseases such as COPD, but lower and less stable performance for rare diseases such as RP.

Jiayi Wang, Jiao Yang, Rui Guo · 0 citations
Open access Aug 2026

Incremental Diagnostic Value of Clinical Information for Large Language Models Across Multiple Organs: Retrospective Study

The addition of clinical information was associated with a numeric trend toward higher diagnostic accuracy overall, but this trend was heterogeneous across models and disease types, and no statistically significant improvement was demonstrated after adjustment for multiple comparisons.

Jin-Qi Zhang, Xiao-Yi Wang, Yanfeng Zhao et al. · 0 citations
Open access Aug 2026

Evaluation of Diagnostic Accuracy of Open-Source and Proprietary Large Language Models Across Multi-System Clinical Cases

A reproducible estimate of diagnostic retrieval accuracy across four widely used model configurations is provided to establish a baseline for further clinical validation and establish a baseline for further clinical validation.

Lalwani Saurabh, Bodetti Dr.Vishala, Gor Kishan et al. · 0 citations
Open access Aug 2026

A Bilingual Benchmark for Evaluating Diagnostic Performance of Multimodal Large Language Models in Radiology (RadM-Bench): Evaluation Development and Validation

Abstract Background Multimodal large language models are increasingly used in radiological diagnosis, but their performance has not been systematically evaluated across volumetric (3D) imaging, real-world clinical versus public teaching cases, and bilingual contexts. Objective The aim of the study is to develop a bilin...

Qing-Xia Wu, Qing-Xia Wu, Pei-Pei Zhang et al. · 0 citations
Open access Aug 2026

Multicenter evaluation of four large language models for automated spine imaging diagnosis

Accurate interpretation of spine imaging is essential for clinical decision-making, yet the diagnostic potential of large language models (LLMs) for radiological report analysis remains inadequately evaluated in terms of sample size, multi-model comparison, reproducibility, and cross-institutional generalisability. Her...

Hao-Lai Liu, Hao Zhang, Hai-Xin Wei et al. · 0 citations
Open access Sep 2026

Exploratory evaluation of large language models in patient-oriented orthopedic MRI report interpretation: a two-stage study

Magnetic resonance imaging (MRI) reports for orthopedic conditions are filled with highly specialized terminology, leading to information asymmetry between clinicians and patients. Deficiencies persist in routine patient education regarding orthopedic imaging results, and a well-documented mismatch exists between...

Quan Zhang, Chen Zhang, Jie Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.