Aug 2026· Seminars in Ophthalmology· pp.
1-7
· 0 citations· 20 references
Medicine
TL;DR
LLM-based chatbots can provide acceptable surgical guidance for straightforward primary surgical cases, but their utility is limited in high-risk or complex clinical settings, and their rationale quality was significantly lower in complex scenarios than in primary ones.
Abstract
Objective
To evaluate the utility, rationality, and safety of glaucoma surgery recommendations generated by three prominent large language models (LLMs) - ChatGPT, Microsoft Copilot, and Google Gemini - when applied to real-world clinical scenarios.
Methods
Retrospective records from a tertiary hospital were converted into standardized scenarios and stratified into "primary" and "complex" glaucoma groups. Each LLM was prompted to suggest a single surgical approach and provide a rationale. A blinded team of glaucoma specialists evaluated the outputs based on six criteria: appropriateness, rationale quality, specificity, adherence to guidelines, feasibility, and safety risk, using a normalized 0-100 scale.
Results
Median overall quality scores across all cases were 80.7 for ChatGPT, 80.7 for Copilot, and 82.7 for Gemini, showing no statistically significant difference in general performance (p = .367). However, case complexity significantly affected performance. For Gemini, appropriateness and rationale quality scores dropped significantly in complex cases and were accompanied by a statistically significant increase in safety risk (p = .009). Although ChatGPT and Copilot demonstrated more stability across groups, their rationale quality was significantly lower in complex scenarios than in primary ones (p = .014 and <0.001, respectively). Pairwise analyses revealed that ChatGPT offered superior rationale quality compared to Copilot, while Gemini exhibited higher specificity.
Conclusions
LLM-based chatbots can provide acceptable surgical guidance for straightforward primary surgical cases, but their utility is limited in high-risk or complex clinical settings. The observed deficiencies in rationale and increased safety risks in complex cases suggest that LLMs should be integrated as auxiliary decision-support tools under expert supervision rather than used as autonomous decision-makers in glaucoma surgery planning.
Large language model-based AI systems produced structured glaucoma-related reasoning with performance that overlapped with attending ophthalmologists but did not establish clinical equivalence, but may have potential as supervised decision-support and educational tools.
Hou-Fa Yin, Lixia Shen, Haiyan Cai et al.· Graefe's archive for clinica...· 0 citations
Current LLMs can match or exceed dermatologists in completeness of guideline-based systemic psoriasis recommendations, particularly in low-risk contexts, however, they remain more prone to critical safety errors in complex, high-stakes scenarios.
K. Xia, L. Min, Dan Jian· Clincal and Experimental Der...· 0 citations
Background: Large language models (LLMs) are increasingly used in patient-facing clinical applications, creating a need for scalable methods to evaluate the accuracy, safety, and appropriateness of their responses. Although physician review remains the reference standard, it is resource-intensive and difficult to scale...
B. Collaço, Nadia G. Wood, Yun-Guo Yu et al.· Bioengineering· 0 citations
Large language models (LLMs), including ChatGPT, Gemini, and DeepSeek, are increasingly used in medicine; however, their performance across multiple clinically relevant domains remains incompletely understood. This cross-sectional comparative study evaluated 225 responses generated by three LLMs to 75 clinical question...
F. Böyük, Aysun Karahan Gün, İsmail Polat Canbolat et al.· Journal of Cardiovascular De...· 0 citations
PURPOSE
Most clinical evaluations of large language models assess factual recall rather than the multi-step reasoning behind operative plans. Reasoning-tuned models, post-trained to generate explicit intermediate reasoning before answering, may better approximate surgical decision-making. We compared two such models fr...
C. Lam, Conor T. Boylan, Adit Ravishankar et al.· Spine Deformity· 0 citations
Large language models (LLMs) may assist in the prevention of diabetes-related foot ulcers; however, their performance in classification may not translate effectively to context-dependent decisions.
This study aimed to evaluate the accuracy, clinical actionability, reproducibility, and safety of five public L...
Yang Wen, Li-Yuan Chen, Xin Deng et al.· Frontiers in Endocrinology· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.