Diagnostic accuracy of multimodal large language models compared with oral medicine specialists: a benchmarking study in salivary gland diseases
Abstract
This study aimed to compare the diagnostic accuracy of multimodal Large Language Models (LLMS), namely, ChatGPT-5, ChatGPT-4o, and Gemini Pro 2.5, with board-certified oral medicine specialists. In this retrospective diagnostic accuracy study, 300 histopathologically confirmed salivary gland disease cases served as the reference standard and were converted into standardized multimodal clinical vignettes balanced by gland type, imaging modality, and difficulty. Three LLMs and an expert panel independently generated ranked differential diagnoses. Accuracy was assessed at Top-1 (primary outcome), Top-3, and Top-5. Agreement was evaluated using Cohen’s κ and Gwet’s AC1. Comparisons employed Cochran’s Q tests with Holm-adjusted McNemar analyses and multivariable logistic regression. At Top-1, ChatGPT-5 achieved 50.3% accuracy, significantly higher than the expert panel (37.3%), corresponding to an absolute difference of 13.0% points (95% CI, 7.9–18.3; p < 0.001), whereas ChatGPT-4o (41.7%) and Gemini Pro 2.5 (42.0%) did not differ significantly from experts. At Top-3, ChatGPT-5 reached 67.3%, exceeding ChatGPT-4o (57.0%) and Gemini Pro 2.5 (57.3%), with no significant difference compared with the expert panel (62.3%). At Top-5, ChatGPT-5 (72.3%) and Gemini Pro 2.5 (72.7%) did not differ significantly from the expert panel (77.7%), whereas ChatGPT-4o showed lower accuracy (67.3%, p < 0.001). The highest agreement was observed for ChatGPT-5 at Top-3 (Gwet’s AC1 = 0.671). In this standardized vignette-based benchmark, ChatGPT-5 demonstrated higher Top-1 diagnostic accuracy than the expert panel while achieving comparable Top-3 and Top-5 performance. However, these findings were obtained under controlled retrospective benchmarking conditions rather than routine clinical practice and therefore require prospective real-world validation before clinical implementation.