Aug 2026· Seminars in ultrasound, CT, and MR· 0 citations· 27 references
Medicine
TL;DR
Publicly available multimodal LLMs showed measurable but heterogeneous performance in grayscale ultrasound-based thyroid nodule classification, but none of the models matched senior radiologist-level performance.
Abstract
Background
Multimodal large language models (LLMs) are increasingly being explored for medical image analysis, but their relative performance in thyroid ultrasound remains unclear.
Objective
This study aimed to compare six publicly available multimodal LLMs for grayscale ultrasound-based classification of thyroid nodules.
Methods
This prospective cross-sectional study included 178 patients with 239 thyroid nodules who underwent preoperative thyroid ultrasound followed by histopathological confirmation. Cropped grayscale ultrasound images of the maximal transverse and longitudinal views were analyzed by six publicly available multimodal LLMs: ChatGPT-5.4, Gemini 3.1 Pro, Claude Opus 4.6, Qwen3.6-Plus, Kimi K2.5, and ERNIE 5.0. All models were evaluated using the same image-input workflow and a standardized prompt, without fine-tuning or task-specific retraining. Agreement was assessed using Cohen's kappa, and diagnostic performance was evaluated using receiver operating characteristic (ROC) analysis. Radiologist benchmarks were included for comparison.
Results
All six LLMs significantly distinguished benign from malignant nodules (all P ≤ 0.001). Gemini 3.1 Pro achieved the best overall performance, with a kappa value of 0.580 and an area under the ROC curve (AUC) of 77.1% (95% CI, 71.5%-82.7%). ChatGPT-5.4 and Qwen3.6-Plus each yielded an AUC of 73.5%, and Kimi K2.5 achieved an AUC of 71.3%. Claude Opus 4.6 and ERNIE 5.0 showed lower overall performance, with AUCs of 65.7% and 59.9%, respectively. The senior radiologist achieved higher diagnostic performance than all six LLMs.
Conclusion
Publicly available multimodal LLMs showed measurable but heterogeneous performance in grayscale ultrasound-based thyroid nodule classification. Gemini 3.1 Pro demonstrated the best overall results, but none of the models matched senior radiologist-level performance.
BACKGROUND
Multimodal large language models (LLMs) have shown potential in medical image analysis, but their performance in the direct interpretation of thyroid ultrasound images remains limited. Whether clinician-recorded structured sonographic features can improve LLM-based thyroid nodule classification is unclear....
This study aimed to compare the diagnostic accuracy of multimodal Large Language Models (LLMS), namely, ChatGPT-5, ChatGPT-4o, and Gemini Pro 2.5, with board-certified oral medicine specialists. In this retrospective diagnostic accuracy study, 300 histopathologically confirmed salivary gland disease cases served as the...
Fatma E. A. Hassanein, Yousra Ahmed, Asmaa Abou-Bakr et al.· The Egyptian Journal of Otol...· 0 citations
Cystoscopic assessment is central to bladder cancer diagnosis, yet visual interpretation remains variable. Existing artificial intelligence approaches often depend on data-intensive models that are often difficult to deploy in routine practice. We evaluated whether multimodal large language models (MLLMs), including sm...
Yonatan Prat, Husny Mahmud, A. Tsur et al.· World journal of urology· 0 citations
Background Pancreatic cystic lesions (PCLs) require precise imaging characterization to guide clinical management. Contrast-enhanced ultrasound (CEUS) reports contain operator-dependent narratives that challenge clinicians. Large language models (LLMs) show potential in medical text analysis but lack validation for pan...
Yu-Su Shao, Yang Gui, Xiao-Yi Yan et al.· Quantitative Imaging in Medi...· 0 citations
Accurate interpretation of spine imaging is essential for clinical decision-making, yet the diagnostic potential of large language models (LLMs) for radiological report analysis remains inadequately evaluated in terms of sample size, multi-model comparison, reproducibility, and cross-institutional generalisability. Her...
Hao-Lai Liu, Hao Zhang, Hai-Xin Wei et al.· npj Digital Medicine· 0 citations
The addition of clinical information was associated with a numeric trend toward higher diagnostic accuracy overall, but this trend was heterogeneous across models and disease types, and no statistically significant improvement was demonstrated after adjustment for multiple comparisons.
Jin-Qi Zhang, Xiao-Yi Wang, Yanfeng Zhao et al.· Journal of Medical Internet...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.