Skip to content
Review Open access

Vision and Language Models for Classifying Maxillary Sinus Disease on Cone-Beam Computed Tomography: A Transparent Multimodal Benchmark

Aug 2026 · medRxiv · 0 citations
Medicine

Abstract

Background: Cone-beam computed tomography (CBCT) frequently captures the maxillary sinuses incidentally, and reliable automated detection of sinus abnormality is clinically relevant. Unlike most vision-language benchmarks in medical imaging, which pair images with pre-existing, human-authored clinical reports, findings text can also be generated directly by a large language model from the image itself--raising the question of how much diagnostic value such AI-derived text carries, and whether that value depends on independent verification. Multimodal artificial intelligence (AI) benchmarks risk overstating performance if the provenance of each input--image, raw AI-generated text, or radiologist-verified text--is not clearly separated and reported. Methods: We used 300 mid-sagittal CBCT slices from the MMDental dataset. ChatGPT generated findings text and a provisional normal/abnormal label for every slice (majority vote, three independent readings from the image alone); primary classification performance was assessed on this full, unfiltered set (n=300). A radiologist then independently reviewed each case's image together with ChatGPT's description, producing their own diagnosis; three cases were excluded as insufficient, yielding 297 confirmed cases. On this subset, every model was retrained and re-evaluated under identical 10-fold cross-validation on both the provisional ChatGPT-only labels ("pre") and the radiologist-confirmed labels ("post"), isolating the effect of label provenance from image or architecture. Eight vision architectures, seven language classifiers, and five VLMs were evaluated throughout; three generative models performed exploratory note-drafting. Findings: Raw ChatGPT-generated text produced the highest performance of any modality or condition: language models reached near-ceiling AUC (0.992 to 1.000, n=300), exceeding every vision model (AUC 0.799 to 0.880) and every VLM image-only probe (AUC 0.63 to 0.69). On the 297-case pre/post analysis, this advantage depended heavily on label source: language and text-derived VLM performance fell substantially from ChatGPT-only to radiologist-confirmed labels (e.g. BERT-base AUC 0.999 to 0.837), while vision-model performance was stable or modestly improved (e.g. DenseNet-121 0.867 to 0.891). The radiologist reclassified 62 of 297 cases (21%) relative to ChatGPT's provisional read, and a meaningful proportion of raw ChatGPT text was clinically uninterpretable or unsupported by the imaging. Interpretation: As shown here for the first time, raw, image-derived AI-generated text yields the highest apparent classification performance in this benchmark, but this reflects the text's alignment with its own self-generated labels rather than verified diagnostic content, and a substantial share of that text is not clinically explainable. Radiologist-confirmed text and labels give a lower but trustworthy estimate of true performance, on which convolutional neural network (CNN) vision models remain a stable, comparatively inexpensive baseline. Multimodal dental AI should report performance separately by modality and label provenance rather than pooling headline metrics.

Read PDF

Similar papers

Open access Aug 2026

Vision-Language Model as a ‘Zero-Shot’ Assistant for Evaluating Condylar Osseous Changes in Cone-beam Computed Tomography

Introduction and aims Interpreting condylar osseous changes on CBCT is challenging for general practitioners. This study evaluated the ‘zero-shot’ diagnostic performance and utility of Vision-Language Models (VLMs) as AI assistants for detecting condylar abnormalities. Methods We analysed 72 CBCT images from the EHPN s...

Ke Chen, Andrew Zhang, Xian-Ju Xie et al. · 0 citations
Open access Sep 2026

Comparison of Multimodal Large Language Models and Oral and Maxillofacial Radiologists in the Detection of Incidental Findings on Panoramic Radiographs: A CBCT-Referenced Diagnostic Accuracy Study

Highlights What are the main findings? In a finding-enriched dataset, two expert oral and maxillofacial radiologists achieved higher overall diagnostic performance than three multimodal (image-capable) LLMs in detecting nine incidental findings on panoramic radiographs against a CBCT-based reference standard (sensitivi...

İsmail Çapar, Utku Cem Hasırcı, Didem Dumanlı Kusay et al. · 0 citations
#small language model Review Open access Sep 2026

Large language models for ophthalmic examination understanding: from information extraction to clinical decision support

Publicly available LLMs and MLLMs are evolving from report-parsing tools toward broader ophthalmic clinical assistants, but their use should follow a task-layered validation framework in which verification and human oversight increase with the clinical consequences of error.

Gang-Yi Wang, Xuan-Qiao Lin, Yi-Zhou Yang · 0 citations
Open access Sep 2026

Comparative evaluation of explainable vision models for bone tumour detection on radiographs

Accurate detection of bone tumours on radiographs can be challenging because lesion appearance varies and interpretation requires specialist expertise. This study evaluates deep learning (DL) models for bone tumour localisation, segmentation, classification, and explainability using the publicly available Bone Tumo...

Daniel Jones Ortega, Khuhed Memon, Syed Saad Azhar Ali et al. · 0 citations
Open access Aug 2026

Pan-retinal pathology detection in oct scans integrating natural language synthesis with diagnostic annotation

iOCT is established as a deployable, multi-disease OCT system with performance approaching trained ophthalmologists, supporting large-scale retinal disease screening in primary care and automated multi-sectional scan analysis and comprehensive natural language report generation.

Wangting Li, Wei-Hao Gao, Lu Chen et al. · 0 citations
#small language model Open access Sep 2026

Multimodal large language models for bladder tumor detection in cystoscopy: a retrospective benchmarking study

Cystoscopic assessment is central to bladder cancer diagnosis, yet visual interpretation remains variable. Existing artificial intelligence approaches often depend on data-intensive models that are often difficult to deploy in routine practice. We evaluated whether multimodal large language models (MLLMs), including sm...

Yonatan Prat, Husny Mahmud, A. Tsur et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.