Diagnostic accuracy of multimodal large language models compared with oral medicine specialists: a benchmarking study in salivary gland diseases
This study aimed to compare the diagnostic accuracy of multimodal Large Language Models (LLMS), namely, ChatGPT-5, ChatGPT-4o, and Gemini Pro 2.5, with board-certified oral medicine specialists. In this retrospective diagnostic accuracy study, 300 histopathologically confirmed salivary gland disease cases served as the...