Comparison of Multimodal Large Language Models and Oral and Maxillofacial Radiologists in the Detection of Incidental Findings on Panoramic Radiographs: A CBCT-Referenced Diagnostic Accuracy Study
Abstract
Highlights What are the main findings? In a finding-enriched dataset, two expert oral and maxillofacial radiologists achieved higher overall diagnostic performance than three multimodal (image-capable) LLMs in detecting nine incidental findings on panoramic radiographs against a CBCT-based reference standard (sensitivity 89.5–91.6% vs. 75.1–83.1%). The gap was widest in the exploratory high-risk group (expert sensitivity above 90% vs. 62.9–78.0% for the LLMs); at the level of individual findings, the difference remained significant after correction for multiple comparisons only for carotid artery calcification and extensive maxillary sinus pathology. What are the implications of the main findings? Under the consumer-interface conditions and access period tested, general-purpose multimodal LLMs should not be relied on independently to evaluate panoramic radiographs for incidental findings, particularly high-risk ones. Any future role for these models is more likely to be as an expert-supervised adjunct or screening aid than as a replacement for expert interpretation; such a combined workflow was not tested here and would require prospective evaluation in clinically representative datasets. Abstract Background/Objectives: This study aimed to compare the diagnostic performance of multimodal (image-capable) large language models (LLMs) and oral and maxillofacial radiologists in detecting nine predefined incidental findings on panoramic radiographs, using a cone-beam computed tomography (CBCT)-derived reference standard. Methods: This retrospective diagnostic performance study included 500 purposively assembled, finding-enriched panoramic radiographs paired with CBCT images. CBCT images were evaluated by three radiologists to establish the reference standard. Two experts and three LLMs, the latter accessed through their consumer web interfaces, independently assessed the presence or absence of the nine findings; 4500 finding-level decisions were analyzed for each reader. Sensitivity, specificity, accuracy, and error rates were calculated. Within-patient clustering was accounted for using generalized estimating equations and a cluster bootstrap procedure with 5000 resamples. Finding-specific comparisons used Cochran’s Q test with Benjamini–Hochberg correction. Results: Agreement between the two experts was very good (κ = 0.86). Expert sensitivity, specificity, and accuracy ranged from 89.5 to 91.6%, 92.0–93.0%, and 91.8–92.9%, respectively, compared with 75.1–83.1%, 88.0–90.0%, and 86.8–89.0% for the LLMs. The overall reader effect was significant for all three performance metrics (all p < 0.001). In the exploratory high-risk group, expert sensitivity ranged from 90.9 to 93.2% versus 62.9–78.0% for the LLMs. After correction for multiple comparisons, the difference between readers remained significant only for carotid artery calcification (q < 0.001) and extensive maxillary sinus pathology (q = 0.005). Conclusions: Under the consumer-interface conditions and access period tested, the LLMs performed below the experts, particularly for high-risk findings, and should not be used independently to evaluate panoramic radiographs. Because the dataset was finding-enriched and single-center, the absolute estimates cannot be transferred directly to routine clinical populations, and any future role for these models, more plausibly as an expert-supervised adjunct or screening aid than as a replacement for expert interpretation, remains to be tested prospectively.