Skip to content

Diagnostic Accuracy of Multimodal Large Language Models in Retinal Fundus Photography.

Aug 2026 · Retina · 0 citations
Medicine

TL;DR

ChatGPT demonstrated the strongest accuracy, justification-accuracy association, and confidence-accuracy correlation compared with Claude, Gemini, and Grok and showed a positive correlation between confidence and accuracy.

Abstract

Purpose

Artificial intelligence in ophthalmology has mostly used task-specific convolutional neural networks. The performance of general-purpose multimodal large language models (LLMs) for fundus image diagnosis remains largely unexplored. This study evaluated and compared the performance of four multimodal LLMs in diagnosing retinal diseases from color fundus photographs.

Methods

A total of 800 images (100 per category across 8 conditions) were randomly sampled and evaluated by ChatGPT-4o, Claude 3.5 Sonnet, Gemini 1.5, and Grok-2. Each model received a standardized prompt with a single fundus photograph and was asked to provide a free-text diagnosis, confidence rating (1-10), and justification type, defined as the reasoning category behind its diagnosis. Primary outcome was overall diagnostic accuracy. Secondary outcomes included sensitivity, specificity, F1-scores, misclassification patterns, justification-accuracy association, confidence-accuracy correlation, Cohen's κ agreement across models, and mean response time per image.

Results

ChatGPT achieved the highest diagnostic accuracy (75%; 603/800), far exceeding Claude (18%), Grok (18%), and Gemini (14%). Unlike other models, ChatGPT demonstrated a meaningful association between justification type and diagnostic correctness and showed a positive correlation between confidence and accuracy (r = 0.19; P < 0.001). Claude and Gemini exhibited weaker confidence-accuracy correlations, while Grok showed none. Agreement across models was consistently low (κ = 0.11-0.16 with ChatGPT; κ ≈ 0.05 among Claude, Gemini, and Grok).

Conclusions

ChatGPT demonstrated the strongest accuracy, justification-accuracy association, and confidence-accuracy correlation compared with Claude, Gemini, and Grok. While promising, refinement, validation, and workflow integration are essential before clinical deployment.

View source

Similar papers

Open access Sep 2026

Comparative Evaluation of EfficientNet-B3, Vision Transformers, and RETFound for Seven-Category Retinopathy Fundus Image Classification

Background/Objectives: Retinal fundus photography is the most widely available means of screening for sight-threatening disease, but it is difficult to read: several conditions leave overlapping signs on the same photograph, the features that separate them are often subtle and confined to small regions of the retina, a...

Majid Alutaibi, Meshari Alazmi · 0 citations
Review Open access Aug 2026

Comparative performance of contemporary multimodal large language models in retinal imaging question answering

These findings support continued benchmarking and cautious, specialist-supervised use of multimodal LLMs in retinal education and structured imaging question answering and indicate that retinal image interpretation remains a major limitation.

Xiao-Chen Gu, Yi-Zhou Yang, Xuan-Qiao Lin et al. · 0 citations
Open access Aug 2026

Pan-retinal pathology detection in oct scans integrating natural language synthesis with diagnostic annotation

While OCT is pivotal for macular disease diagnosis, its adoption in primary care is limited by AI systems that cannot simultaneously analyze multi-sectional scans across the full spectrum of maculopathies or generate diagnostically integrated reports. Here we present iOCT, an intelligent OCT analysis system that in...

Wangting Li, Wei-Hao Gao, Lu Chen et al. · 0 citations
Preprint Aug 2026

An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography

Large language models (LLMs) show promise in medical image interpretation but suffer from hallucination, limited accuracy, and run-to-run inconsistency. We developed and validated an agentic AI framework integrating LLMs with specialized deep learning tools for glaucoma detection from fundus photography. The workflow h...

J. Jalili, H. Taghizad, Anuwat Jiravarnsirikul et al. · 0 citations
Open access Sep 2026

Robust and interpretable retinal disease classification: a cross-source validation study

Retinal diseases such as hypertensive retinopathy (HR), myopia, and retinitis pigmentosa (RP) may remain undiagnosed until irreversible visual impairment occurs. Retinal multi disease classification models are often developed by integrating multiple public fundus datasets and evaluated using internal validation. We e...

A. Manal, Fathima Hafsa, Z. Khalid et al. · 0 citations
Conference Aug 2026

Multi-Disease Classification in Retinal Oct Using Deep Transfer Learning: A Comparative Study

Timely and precise detection of retinal diseases is essential for preventing permanent vision loss. However, the analysis of Optical Coherence Tomography (OCT) images is often time-taking and subject to inter-clinician variability. This paper presents deep learning architectures for multi-disease classification of OCT...

V. A. Balakrishna Jakka, P. Naganjaneyulu, Narra Dhana Lakshmi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.