Skip to content
#small language model Review Open access

Large language models for ophthalmic examination understanding: from information extraction to clinical decision support

Sep 2026 · Frontiers in Medicine · Vol 13 · 0 citations · 66 references
Medicine

TL;DR

Publicly available LLMs and MLLMs are evolving from report-parsing tools toward broader ophthalmic clinical assistants, but their use should follow a task-layered validation framework in which verification and human oversight increase with the clinical consequences of error.

Abstract

Ophthalmology relies on heterogeneous examination outputs, including device-generated reports, structured measurements, fundus photographs, optical coherence tomography (OCT), visual field plots, corneal imaging, and multimodal clinical data. Publicly available large language models (LLMs) and multimodal large language models (MLLMs) can process these inputs without study-specific ophthalmic training, creating opportunities for information extraction, interpretation, and clinical decision support. This Mini Review synthesizes 37 peer-reviewed studies evaluating such models across ophthalmic examination types and task layers. The evidence shows a consistent task gradient. Structured extraction and constrained calculations are generally more reliable than open-ended image interpretation, multimodal diagnosis, or treatment planning. Performance improves when inputs are standardized, clinical context is provided, and prompts or decision rules are constrained, but remains sensitive to report layout, image preprocessing, prompt design, and model updates. Common limitations include retrospective or public datasets, small or selected cohorts, limited external validation, possible training-data contamination, inconsistent reporting of prompts and model access, and a focus on technical accuracy rather than clinical outcomes. Clinical translation therefore requires safeguards matched to task risk: structured outputs and deterministic checks for extraction, traceable evidence and cross-device validation for interpretation, and independent verification, deferral mechanisms, and clinician oversight for decision support. Publicly available LLMs and MLLMs are evolving from report-parsing tools toward broader ophthalmic clinical assistants, but their use should follow a task-layered validation framework in which verification and human oversight increase with the clinical consequences of error.

Read PDF

Similar papers

Review Open access Aug 2026

Vision and Language Models for Classifying Maxillary Sinus Disease on Cone-Beam Computed Tomography: A Transparent Multimodal Benchmark

Background: Cone-beam computed tomography (CBCT) frequently captures the maxillary sinuses incidentally, and reliable automated detection of sinus abnormality is clinically relevant. Unlike most vision-language benchmarks in medical imaging, which pair images with pre-existing, human-authored clinical reports, findings...

S. Alhebshi, H. Khalifa, T. D. Pham · 0 citations
#artificial intelligence Preprint Sep 2026

EyeVQA: Benchmarking Ophthalmic Vision-Language Models from Recognition to Spatial Grounding

Vision-language models (VLMs) have shown increasing potential for medical image understanding, yet their capabilities in ophthalmic imaging remain insufficiently characterized. Existing ophthalmic datasets are typically designed for individual diseases or specialized tasks, making it difficult to systematically evaluat...

Gu-Jie Shao, Zi-Xun Xie, Xue-Chun Xing et al. · 0 citations
Aug 2026

Diagnostic Accuracy of Multimodal Large Language Models in Retinal Fundus Photography.

ChatGPT demonstrated the strongest accuracy, justification-accuracy association, and confidence-accuracy correlation compared with Claude, Gemini, and Grok and showed a positive correlation between confidence and accuracy.

Lia Huo, Astha Chandra, Michael Balas et al. · 0 citations
Aug 2026

Comparative evaluation of large language models and clinicians in real-world glaucoma clinical reasoning.

Large language model-based AI systems produced structured glaucoma-related reasoning with performance that overlapped with attending ophthalmologists but did not establish clinical equivalence, but may have potential as supervised decision-support and educational tools.

Hou-Fa Yin, Lixia Shen, Haiyan Cai et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 30, 2026

This game-playing AI is the new champ at Stratego

Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.