Skip to content

Comparative Performance of Multimodal Large Language Models in Grayscale Ultrasound-Based Classification of Thyroid Nodules.

Aug 2026 · Seminars in ultrasound, CT, and MR · 0 citations · 27 references
Medicine

TL;DR

Publicly available multimodal LLMs showed measurable but heterogeneous performance in grayscale ultrasound-based thyroid nodule classification, but none of the models matched senior radiologist-level performance.

Abstract

Background

Multimodal large language models (LLMs) are increasingly being explored for medical image analysis, but their relative performance in thyroid ultrasound remains unclear.

Objective

This study aimed to compare six publicly available multimodal LLMs for grayscale ultrasound-based classification of thyroid nodules.

Methods

This prospective cross-sectional study included 178 patients with 239 thyroid nodules who underwent preoperative thyroid ultrasound followed by histopathological confirmation. Cropped grayscale ultrasound images of the maximal transverse and longitudinal views were analyzed by six publicly available multimodal LLMs: ChatGPT-5.4, Gemini 3.1 Pro, Claude Opus 4.6, Qwen3.6-Plus, Kimi K2.5, and ERNIE 5.0. All models were evaluated using the same image-input workflow and a standardized prompt, without fine-tuning or task-specific retraining. Agreement was assessed using Cohen's kappa, and diagnostic performance was evaluated using receiver operating characteristic (ROC) analysis. Radiologist benchmarks were included for comparison.

Results

All six LLMs significantly distinguished benign from malignant nodules (all P ≤ 0.001). Gemini 3.1 Pro achieved the best overall performance, with a kappa value of 0.580 and an area under the ROC curve (AUC) of 77.1% (95% CI, 71.5%-82.7%). ChatGPT-5.4 and Qwen3.6-Plus each yielded an AUC of 73.5%, and Kimi K2.5 achieved an AUC of 71.3%. Claude Opus 4.6 and ERNIE 5.0 showed lower overall performance, with AUCs of 65.7% and 59.9%, respectively. The senior radiologist achieved higher diagnostic performance than all six LLMs.

Conclusion

Publicly available multimodal LLMs showed measurable but heterogeneous performance in grayscale ultrasound-based thyroid nodule classification. Gemini 3.1 Pro demonstrated the best overall results, but none of the models matched senior radiologist-level performance.

View source

Similar papers

Sep 2026

Impact of Clinician-Recorded ACR TI-RADS Features on ChatGPT-5.4-Based Thyroid Nodule Classification.

BACKGROUND Multimodal large language models (LLMs) have shown potential in medical image analysis, but their performance in the direct interpretation of thyroid ultrasound images remains limited. Whether clinician-recorded structured sonographic features can improve LLM-based thyroid nodule classification is unclear....

Zi-Man Chen, Fei Chen, Ying-Li Wang · 0 citations
Open access Sep 2026

Diagnostic accuracy of multimodal large language models compared with oral medicine specialists: a benchmarking study in salivary gland diseases

This study aimed to compare the diagnostic accuracy of multimodal Large Language Models (LLMS), namely, ChatGPT-5, ChatGPT-4o, and Gemini Pro 2.5, with board-certified oral medicine specialists. In this retrospective diagnostic accuracy study, 300 histopathologically confirmed salivary gland disease cases served as the...

Fatma E. A. Hassanein, Yousra Ahmed, Asmaa Abou-Bakr et al. · 0 citations
#small language model Open access Sep 2026

Multimodal large language models for bladder tumor detection in cystoscopy: a retrospective benchmarking study

Cystoscopic assessment is central to bladder cancer diagnosis, yet visual interpretation remains variable. Existing artificial intelligence approaches often depend on data-intensive models that are often difficult to deploy in routine practice. We evaluated whether multimodal large language models (MLLMs), including sm...

Yonatan Prat, Husny Mahmud, A. Tsur et al. · 0 citations
Open access Aug 2026

Large language models for analyzing contrast-enhanced ultrasound reports of pancreatic cystic lesions

Background Pancreatic cystic lesions (PCLs) require precise imaging characterization to guide clinical management. Contrast-enhanced ultrasound (CEUS) reports contain operator-dependent narratives that challenge clinicians. Large language models (LLMs) show potential in medical text analysis but lack validation for pan...

Yu-Su Shao, Yang Gui, Xiao-Yi Yan et al. · 0 citations
Open access Aug 2026

Multicenter evaluation of four large language models for automated spine imaging diagnosis

Accurate interpretation of spine imaging is essential for clinical decision-making, yet the diagnostic potential of large language models (LLMs) for radiological report analysis remains inadequately evaluated in terms of sample size, multi-model comparison, reproducibility, and cross-institutional generalisability. Her...

Hao-Lai Liu, Hao Zhang, Hai-Xin Wei et al. · 0 citations
Open access Aug 2026

Incremental Diagnostic Value of Clinical Information for Large Language Models Across Multiple Organs: Retrospective Study

The addition of clinical information was associated with a numeric trend toward higher diagnostic accuracy overall, but this trend was heterogeneous across models and disease types, and no statistically significant improvement was demonstrated after adjustment for multiple comparisons.

Jin-Qi Zhang, Xiao-Yi Wang, Yanfeng Zhao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.