Skip to content
Review Open access

Comparative performance of contemporary multimodal large language models in retinal imaging question answering

Aug 2026 · Frontiers in Medicine · Vol 13 · 0 citations · 33 references
Medicine

TL;DR

These findings support continued benchmarking and cautious, specialist-supervised use of multimodal LLMs in retinal education and structured imaging question answering and indicate that retinal image interpretation remains a major limitation.

Abstract

Background Contemporary multimodal large language models (LLMs) can process both clinical text and medical images, but their reliability in retinal imaging-based question answering remains uncertain. This study evaluated the performance of six contemporary multimodal LLMs on retinal multiple-choice questions derived from OCTCases. Methods This cross-sectional benchmark study included 226 multiple-choice questions from 78 OCTCases retinal cases, comprising 151 image-based and 75 nonimage-based questions. Six multimodal LLMs were evaluated: ChatGPT-5.5 Instant, ChatGPT-5.5 Thinking, Grok 4, Gemini 3, DeepSeek V4, and Kimi 2.5. Model-selected answers were compared with the OCTCases answer key, which was reviewed by three retina specialists. Accuracy was assessed overall and by question subset. Pairwise model comparisons, inter-model agreement, voting-based group-answer performance, item difficulty distribution, and interface latency were analyzed. Results Overall accuracy ranged from 59.3 to 77.4%, with the highest accuracy observed for ChatGPT-5.5 Thinking, followed by Gemini 3 and ChatGPT-5.5 Instant. DeepSeek V4 showed the lowest overall accuracy. All models performed better on nonimage-based questions than on image-based questions. In the image-based subset, Gemini 3 achieved the highest accuracy, whereas DeepSeek V4 showed the lowest accuracy. Under the majority-vote rule, group answers were generated for 182 of 226 questions and achieved an accuracy of 89.6% among questions with a determinate group answer. The plurality-vote rule generated group answers for more questions but with lower accuracy. Inter-model agreement was generally higher for nonimage-based than image-based questions, and all questions answered incorrectly by all six models were image-based. Interface latency varied across models and showed no consistent association with response correctness. Conclusion Contemporary multimodal LLMs showed promising but uneven performance on OCTCases retinal questions. Their performance was consistently better on nonimage-based than image-based questions, indicating that retinal image interpretation remains a major limitation. Voting-based group answers may provide a useful reliability signal when model agreement is present, whereas model disagreement may help identify difficult or visually ambiguous questions requiring specialist review. These findings support continued benchmarking and cautious, specialist-supervised use of multimodal LLMs in retinal education and structured imaging question answering.

Read PDF

Similar papers

Aug 2026

Diagnostic Accuracy of Multimodal Large Language Models in Retinal Fundus Photography.

ChatGPT demonstrated the strongest accuracy, justification-accuracy association, and confidence-accuracy correlation compared with Claude, Gemini, and Grok and showed a positive correlation between confidence and accuracy.

Lia Huo, Astha Chandra, Michael Balas et al. · 0 citations
Aug 2026

Comparative Performance of Multimodal Large Language Models in Grayscale Ultrasound-Based Classification of Thyroid Nodules.

Publicly available multimodal LLMs showed measurable but heterogeneous performance in grayscale ultrasound-based thyroid nodule classification, but none of the models matched senior radiologist-level performance.

Zi-Man Chen, Ying-Li Wang, Fei Chen · 0 citations
Open access Sep 2026

Performance of vision-language models compared with 252 medical students on text-only and image-based dermatology examinations

Vision–language models (VLMs) are increasingly evaluated in medical education, yet their performance on visually intensive assessments remains incompletely understood. We compared four state-of-the-art VLMs, GPT-4o and GPT-5 (both accessed via ChatGPT), Gemini 2.5 Flash, and Gemini 3 Pro, with fifth-year medical studen...

O. Erdem, A. Yilmaz, A. Şahin et al. · 0 citations
Open access Sep 2026

Do Large Language Models Use the Clinical Vignette? A Question Ablation Study on the Orthopaedic In-Training Examination

Background: Large language models (LLMs) have demonstrated strong performance on standardized medical examinations, with recent studies reporting performance approaching or exceeding that of senior medical residents. However, examination accuracy alone does not establish how models arrive at their answers or the relati...

F. Gafoor, M. Syed, M. Halai et al. · 0 citations
#artificial intelligence Preprint Sep 2026

EyeVQA: Benchmarking Ophthalmic Vision-Language Models from Recognition to Spatial Grounding

Vision-language models (VLMs) have shown increasing potential for medical image understanding, yet their capabilities in ophthalmic imaging remain insufficiently characterized. Existing ophthalmic datasets are typically designed for individual diseases or specialized tasks, making it difficult to systematically evaluat...

Gu-Jie Shao, Zi-Xun Xie, Xue-Chun Xing et al. · 0 citations
Aug 2026

Do Accuracy Gains Reflect Genuine Visual Understanding? A Multi-Model Evaluation of Vision-Language Models in Thoracic Imaging.

Even at the highest accuracy observed here, answer-level performance can overstate visual understanding, pointing to the need for more appropriate evaluation methods that assess whether a model's reasoning is faithful to the image, not only whether its answer is correct.

Jiyoung Song, Won Gi Jeong, Hongseok Ko et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.