ChatGPT demonstrated the strongest accuracy, justification-accuracy association, and confidence-accuracy correlation compared with Claude, Gemini, and Grok and showed a positive correlation between confidence and accuracy.
Abstract
Purpose
Artificial intelligence in ophthalmology has mostly used task-specific convolutional neural networks. The performance of general-purpose multimodal large language models (LLMs) for fundus image diagnosis remains largely unexplored. This study evaluated and compared the performance of four multimodal LLMs in diagnosing retinal diseases from color fundus photographs.
Methods
A total of 800 images (100 per category across 8 conditions) were randomly sampled and evaluated by ChatGPT-4o, Claude 3.5 Sonnet, Gemini 1.5, and Grok-2. Each model received a standardized prompt with a single fundus photograph and was asked to provide a free-text diagnosis, confidence rating (1-10), and justification type, defined as the reasoning category behind its diagnosis. Primary outcome was overall diagnostic accuracy. Secondary outcomes included sensitivity, specificity, F1-scores, misclassification patterns, justification-accuracy association, confidence-accuracy correlation, Cohen's κ agreement across models, and mean response time per image.
Results
ChatGPT achieved the highest diagnostic accuracy (75%; 603/800), far exceeding Claude (18%), Grok (18%), and Gemini (14%). Unlike other models, ChatGPT demonstrated a meaningful association between justification type and diagnostic correctness and showed a positive correlation between confidence and accuracy (r = 0.19; P < 0.001). Claude and Gemini exhibited weaker confidence-accuracy correlations, while Grok showed none. Agreement across models was consistently low (κ = 0.11-0.16 with ChatGPT; κ ≈ 0.05 among Claude, Gemini, and Grok).
Conclusions
ChatGPT demonstrated the strongest accuracy, justification-accuracy association, and confidence-accuracy correlation compared with Claude, Gemini, and Grok. While promising, refinement, validation, and workflow integration are essential before clinical deployment.
Background/Objectives: Retinal fundus photography is the most widely available means of screening for sight-threatening disease, but it is difficult to read: several conditions leave overlapping signs on the same photograph, the features that separate them are often subtle and confined to small regions of the retina, a...
These findings support continued benchmarking and cautious, specialist-supervised use of multimodal LLMs in retinal education and structured imaging question answering and indicate that retinal image interpretation remains a major limitation.
Xiao-Chen Gu, Yi-Zhou Yang, Xuan-Qiao Lin et al.· Frontiers in Medicine· 0 citations
While OCT is pivotal for macular disease diagnosis, its adoption in primary care is limited by AI systems that cannot simultaneously analyze multi-sectional scans across the full spectrum of maculopathies or generate diagnostically integrated reports. Here we present iOCT, an intelligent OCT analysis system that in...
Wangting Li, Wei-Hao Gao, Lu Chen et al.· npj Digital Medicine· 0 citations
Large language models (LLMs) show promise in medical image interpretation but suffer from hallucination, limited accuracy, and run-to-run inconsistency. We developed and validated an agentic AI framework integrating LLMs with specialized deep learning tools for glaucoma detection from fundus photography. The workflow h...
J. Jalili, H. Taghizad, Anuwat Jiravarnsirikul et al.· 0 citations
Retinal diseases such as hypertensive retinopathy (HR), myopia, and retinitis pigmentosa (RP) may remain undiagnosed until irreversible visual impairment occurs. Retinal multi disease classification models are often developed by integrating multiple public fundus datasets and evaluated using internal validation. We e...
A. Manal, Fathima Hafsa, Z. Khalid et al.· Scientific Reports· 0 citations
Timely and precise detection of retinal diseases is essential for preventing permanent vision loss. However, the analysis of Optical Coherence Tomography (OCT) images is often time-taking and subject to inter-clinician variability. This paper presents deep learning architectures for multi-disease classification of OCT...
V. A. Balakrishna Jakka, P. Naganjaneyulu, Narra Dhana Lakshmi et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.