A reasoning-driven vision-language framework that explicitly models the ophthalmologist's diagnostic workflow by generating structured clinical reasoning prior to diagnosis is developed, demonstrating that explicitly modeling expert clinical reasoning simultaneously improves interpretability and diagnostic performance.
Abstract
Glaucoma is a leading cause of irreversible blindness worldwide. Ophthalmologists diagnose glaucoma through a structured reasoning process by sequentially evaluating optic nerve head characteristics before reaching a final diagnosis, whereas existing AI systems typically perform direct image classification without providing clinically meaningful reasoning. We present the first clinically annotated fundus reasoning dataset, comprising 1,077 fundus photographs paired with expert-authored six-step diagnostic reports. Building on this dataset, we develop a reasoning-driven vision-language framework that explicitly models the ophthalmologist's diagnostic workflow by generating structured clinical reasoning prior to diagnosis. The generated reports are clinically validated, achieving the best performance across all evaluated clinical findings, including a cup-to-disc ratio mean absolute error of 0.070, an ISNT Kendall distance of 1.73, and the highest semantic agreement with expert reports (BERTScore-F1 = 0.874). The resulting framework also improves glaucoma diagnosis, achieving a balanced accuracy of $94.7%$ and precision of $94.8%$, demonstrating that explicitly modeling expert clinical reasoning simultaneously improves interpretability and diagnostic performance. Code and data are available at url{https://glaucoma-cot.github.io/}.
Large language model-based AI systems produced structured glaucoma-related reasoning with performance that overlapped with attending ophthalmologists but did not establish clinical equivalence, but may have potential as supervised decision-support and educational tools.
Hou-Fa Yin, Lixia Shen, Haiyan Cai et al.· Graefe's archive for clinica...· 0 citations
ChatGPT demonstrated the strongest accuracy, justification-accuracy association, and confidence-accuracy correlation compared with Claude, Gemini, and Grok and showed a positive correlation between confidence and accuracy.
Lia Huo, Astha Chandra, Michael Balas et al.· Retina· 0 citations
Glaucoma is a chronic, progressive optic neuropathy and a leading cause of irreversible blindness worldwide. Early detection is crucial but is often limited by nonspecific clinical signs and the need for specialist expertise. Artificial intelligence (AI), particularly deep learning, has shown promise for glaucoma detec...
N. Phan, Trong Van Pham, Kim Thanh Doan et al.· JOURNAL OF CURRENT SCIENCE A...· 0 citations
Diabetic Retinopathy (DR) is a leading cause of preventable blindness, requiring accurate and timely diagnosis for effective treatment. While deep learning models achieve high classification accuracy, they often lack clinical interpretability. Conversely, large language models (LLMs) provide strong reasoning capabiliti...
Sait Suer, R. Aygun, Mahmut Karakaya et al.· 2026 International Conferenc...· 0 citations
Large language models (LLMs) show promise in medical image interpretation but suffer from hallucination, limited accuracy, and run-to-run inconsistency. We developed and validated an agentic AI framework integrating LLMs with specialized deep learning tools for glaucoma detection from fundus photography. The workflow h...
J. Jalili, H. Taghizad, Anuwat Jiravarnsirikul et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.