Jan 2026· Journal of Ophthalmology· Vol 2026· 0 citations· 30 references
Medicine
TL;DR
Both models aligned closely with the guidelines, with significant concordance in GRADE‐based recommendation strength, and may serve as preclinical reference tools in refractive surgery but not as a substitute for specialist judgment.
Abstract
Purpose To evaluate the guideline knowledge alignment of two large language models (LLMs), GPT‐5.5 Instant and DeepSeek‐V4, and to determine their preclinical reliability as reference tools in refractive surgery. Methods Using the 38 evidence‐based recommendations of the international keratorefractive lenticule extraction (KLEx) guidelines as the gold standard, both LLMs were evaluated in their default configurations. Two ophthalmologists independently assessed the clinical safety and medical accuracy of the model responses using a 5‐point Likert scale (1–5 points). Agreement between each model’s recommendation strength and the guideline was quantified by intraclass correlation coefficient (ICC), structural reliability by the DISCERN scale, and readability by the Flesch Reading Ease (FRE) and Flesch–Kincaid Grade Level (FKGL) indices. Results Likert ratings did not differ between GPT‐5.5 Instant (4.96 ± 0.206) and DeepSeek‐V4 (4.89 ± 0.385; p = 0.134). Across 114 independent generations, the ICC for agreement with the guideline was 0.884 (95% CI, 0.836–0.918) for GPT‐5.5 Instant and 0.739 (0.640–0.813) for DeepSeek‐V4 (both p < 0.001). DISCERN scores were 70.18 ± 5.16 and 68.05 ± 5.41 (p = 0.085); FRE, 9.93 ± 8.33 and 3.74 ± 5.24 (p < 0.001); and FKGL, 16.07 ± 2.21 and 18.95 ± 1.96 (p < 0.001). Conclusion Both models aligned closely with the guidelines, with significant concordance in GRADE‐based recommendation strength. They may serve as preclinical reference tools in refractive surgery but not as a substitute for specialist judgment.
LLMs demonstrated similar guideline concordance, suggesting patients can expect comparable accuracy across platforms, and LLMs generally improved FKGL scores compared to the AAO-HNS CPG, they demonstrated lower FRE scores, indicating mixed results on overall readability.
Hetal Lad, Emily S Kwon, Ayushi Chadha et al.· Journal of Otorhinolaryngolo...· 0 citations
How well large language models (LLM) handle Turkish melanoma patient education varies widely from one model to the next, and findings suggest that LLM-generated Turkish melanoma materials may be useful as preliminary educational drafts.
Nıyazı Çetın, A. Atılan· European Journal of Therapeu...· 0 citations
This study provides a benchmark of LLM performance for radiology OSCE-style content generation and evaluation during a specific snapshot of artificial intelligence development (August 2024).
Ankush Ankush, Samriddhi Burman, Sydney Smith et al.· Radiology Advances· 0 citations
AI models demonstrate generally high accuracy in medical examinations; however, they struggle with interpreting clinical context, recognising atypical medical scenarios and answering questions involving visual content.
Bünyamin Arı· Health Informatics Journal· 0 citations
There may be significant differences in how effectively LLMs support patients with surgical queries, particularly in areas needing detailed explanation, and usually required minimal clarification in areas needing detailed explanation.
T. Davis, B. Guevel, K. Logishetty et al.· Annals of the Royal College...· 0 citations
It is suggested that model selection influences expert-rated educational suitability and overall information quality, whereas topic complexity primarily affects readability.
Hao Wei, Sisi Sun, Ming-Xin Liu et al.· Journal of Visualized Experi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.