Aug 2026· Bolu Abant Izzet Baysal Universitesi, Tip Fakultesi, Abant Tip Dergisi· 0 citations· 17 references
TL;DR
The evaluated LLMs showed high concordance with EAU 2026 first-line treatment recommendations for ureteral stones, and Concordance was complete in non–priority-sensitive scenarios requiring URS prioritization.
Abstract
ObjectiveLarge language models (LLMs) are increasingly used to answer medical questions, but their reliability in guideline-based urological decisions remains uncertain. This study aimed to evaluate the concordance of three widely used LLMs with the European Association of Urology (EAU) 2026 guideline recommendations for the active management of ureteral stones.Materials and MethodsIn this cross-sectional, vignette-based comparative study, the EAU 2026 ureteral-stone treatment algorithm was converted into 40 standardized clinical vignettes (four groups of ten: proximal 10 mm, distal 10 mm). ChatGPT, Gemini, and Claude were queried with the same standardized prompt, which did not name a specific guideline. Responses were scored against predefined EAU-based reference answers using a binary system. Concordance was compared with Cochran’s Q test.ResultsA total of 120 LLM-generated responses were evaluated. Overall concordance was 96.7% (116/120). ChatGPT achieved complete concordance (40/40, 100%), while Gemini and Claude each achieved 95% (38/40); the difference was not significant (Cochran’s Q=2.67, p=0.264). Concordance was complete in all 10 mm scenarios requiring URS prioritization. All four discordances were “incorrect prioritization” in >10 mm stones, presenting shock-wave lithotripsy as co-equal to ureteroscopy; each involved cross-guideline conflation with American Urological Association (AUA) framing. No unsafe recommendation or guideline hallucination was observed under the predefined scoring categories.ConclusionThe evaluated LLMs showed high concordance with EAU 2026 first-line treatment recommendations for ureteral stones. Concordance was complete in non–priority-sensitive
Final recommendation concordance did not meet the protocol-defined benchmark, ChatGPT-4o never altered an MTB decision, and clinically relevant errors occurred even among highly concordant outputs, showing that concordance alone does not guarantee safety.
J. De la Torre-Trillo, Albert Munuera, M. D. Ureña et al.· Clinical and Translational O...· 0 citations
PURPOSE
To compare the concordance of ChatGPT, Gemini, and Claude with a prespecified expert guideline-based reference standard in fabricated endometrial cancer clinical vignettes under standardized prompting.
METHODS
We conducted a case-based in silico benchmarking study using 35 fabricated postoperative endometrial...
E. Perrone, Giuseppe Parisi, M. Giuliano et al.· JCO Clinical Cancer Informat...· 0 citations
LLMs demonstrated similar guideline concordance, suggesting patients can expect comparable accuracy across platforms, and LLMs generally improved FKGL scores compared to the AAO-HNS CPG, they demonstrated lower FRE scores, indicating mixed results on overall readability.
Hetal Lad, Emily S Kwon, Ayushi Chadha et al.· Journal of Otorhinolaryngolo...· 0 citations
RATIONALE AND OBJECTIVES
Large language models may help radiologists apply increasingly complex imaging guidelines, but their safety and clinical usefulness as decision-support tools in abdominal radiology remain uncertain. We evaluated the usefulness, potential harm, guideline concordance, and triage performance of Ge...
Mehmet Fatih Kaya, Soheil Sabet· European Journal of Radiolog...· 0 citations
Findings support the use of guideline grounding to enhance the guideline concordance and reliability of LLM-generated responses while emphasizing the continued need for clinician oversight.
Sena Kaşıkçı, Ebru Şirinoğlu, Olcay Özdemir· Odontology : official journa...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.