Skip to content

Evaluation of Large Language Model-Generated Recommendations in Glaucoma Surgical Decision-Making.

Aug 2026 · Seminars in Ophthalmology · pp. 1-7 · 0 citations · 20 references
Medicine

TL;DR

LLM-based chatbots can provide acceptable surgical guidance for straightforward primary surgical cases, but their utility is limited in high-risk or complex clinical settings, and their rationale quality was significantly lower in complex scenarios than in primary ones.

Abstract

Objective

To evaluate the utility, rationality, and safety of glaucoma surgery recommendations generated by three prominent large language models (LLMs) - ChatGPT, Microsoft Copilot, and Google Gemini - when applied to real-world clinical scenarios.

Methods

Retrospective records from a tertiary hospital were converted into standardized scenarios and stratified into "primary" and "complex" glaucoma groups. Each LLM was prompted to suggest a single surgical approach and provide a rationale. A blinded team of glaucoma specialists evaluated the outputs based on six criteria: appropriateness, rationale quality, specificity, adherence to guidelines, feasibility, and safety risk, using a normalized 0-100 scale.

Results

Median overall quality scores across all cases were 80.7 for ChatGPT, 80.7 for Copilot, and 82.7 for Gemini, showing no statistically significant difference in general performance (p = .367). However, case complexity significantly affected performance. For Gemini, appropriateness and rationale quality scores dropped significantly in complex cases and were accompanied by a statistically significant increase in safety risk (p = .009). Although ChatGPT and Copilot demonstrated more stability across groups, their rationale quality was significantly lower in complex scenarios than in primary ones (p = .014 and <0.001, respectively). Pairwise analyses revealed that ChatGPT offered superior rationale quality compared to Copilot, while Gemini exhibited higher specificity.

Conclusions

LLM-based chatbots can provide acceptable surgical guidance for straightforward primary surgical cases, but their utility is limited in high-risk or complex clinical settings. The observed deficiencies in rationale and increased safety risks in complex cases suggest that LLMs should be integrated as auxiliary decision-support tools under expert supervision rather than used as autonomous decision-makers in glaucoma surgery planning.

View source

Similar papers

Aug 2026

Comparative evaluation of large language models and clinicians in real-world glaucoma clinical reasoning.

Large language model-based AI systems produced structured glaucoma-related reasoning with performance that overlapped with attending ophthalmologists but did not establish clinical equivalence, but may have potential as supervised decision-support and educational tools.

Hou-Fa Yin, Lixia Shen, Haiyan Cai et al. · 0 citations

Guideline-Based Evaluation of Large Language Models in Psoriasis Treatment.

Current LLMs can match or exceed dermatologists in completeness of guideline-based systemic psoriasis recommendations, particularly in low-risk contexts, however, they remain more prone to critical safety errors in complex, high-stakes scenarios.

K. Xia, L. Min, Dan Jian · 0 citations
Review Open access Sep 2026

Development and Clinical Validation of an Automated LLM Judge for Evaluating Perioperative Patient Questions

Background: Large language models (LLMs) are increasingly used in patient-facing clinical applications, creating a need for scalable methods to evaluate the accuracy, safety, and appropriateness of their responses. Although physician review remains the reference standard, it is resource-intensive and difficult to scale...

B. Collaço, Nadia G. Wood, Yun-Guo Yu et al. · 0 citations
Open access Sep 2026

A Multidimensional Evaluation of Large Language Model Responses to the 2024 ESC Hypertension Guidelines: A Comparative Study

Large language models (LLMs), including ChatGPT, Gemini, and DeepSeek, are increasingly used in medicine; however, their performance across multiple clinically relevant domains remains incompletely understood. This cross-sectional comparative study evaluated 225 responses generated by three LLMs to 75 clinical question...

F. Böyük, Aysun Karahan Gün, İsmail Polat Canbolat et al. · 0 citations
Review Open access Sep 2026

Evaluating reasoning-tuned large language models for clinical decision-making in spine surgery.

PURPOSE Most clinical evaluations of large language models assess factual recall rather than the multi-step reasoning behind operative plans. Reasoning-tuned models, post-trained to generate explicit intermediate reasoning before answering, may better approximate surgical decision-making. We compared two such models fr...

C. Lam, Conor T. Boylan, Adit Ravishankar et al. · 0 citations
Review Open access Sep 2026

From risk classification to clinical action: a prespecified paired pilot benchmark of public large language model interfaces for diabetes-related foot ulcer prevention

Large language models (LLMs) may assist in the prevention of diabetes-related foot ulcers; however, their performance in classification may not translate effectively to context-dependent decisions. This study aimed to evaluate the accuracy, clinical actionability, reproducibility, and safety of five public L...

Yang Wen, Li-Yuan Chen, Xin Deng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.