Skip to content

Accuracy and Safety of Large Language Models in Endometrial Cancer Decision Making: A Case-Based In Silico Benchmarking Study.

Jul 2026 · JCO Clinical Cancer Informatics · Vol 10 3, pp. e2600142 · 0 citations · 17 references
Medicine

Abstract

Purpose

To compare the concordance of ChatGPT, Gemini, and Claude with a prespecified expert guideline-based reference standard in fabricated endometrial cancer clinical vignettes under standardized prompting.

Methods

We conducted a case-based in silico benchmarking study using 35 fabricated postoperative endometrial cancer vignettes representing a broad spectrum of ESGO-ESTRO-ESP 2025 management scenarios. Each vignette was submitted to ChatGPT, Gemini, and Claude in independent chat sessions using the same standardized prompt. The primary end point was concordance with the prespecified expert reference standard, scored as 0 (discordant), 1 (partially concordant), or 2 (fully concordant). Secondary end points were major safety issues and recognition of missing critical information. A post hoc subgroup analysis evaluated the effect of guideline-informed prompting in 12 cases.

Results

Concordance differed significantly across models (P < .001). Gemini achieved the highest performance, with a median concordance score of 2 (IQR 1-2), compared with 1 (IQR 0-1) for ChatGPT and 0 (IQR 0-1) for Claude. Fully concordant recommendations were generated in 65.7% of cases by Gemini, 2.9% by ChatGPT, and 8.6% by Claude. Major safety issues also differed across models (P = .007), occurring in 22.9% of Gemini responses, 48.6% of ChatGPT responses, and 54.3% of Claude responses. In the subset of vignettes with intentionally missing decisive information, Gemini identified the need for additional data in 83.3% of cases, compared with 50.0% for both ChatGPT and Claude. In the post hoc subgroup analysis, guideline-informed prompting significantly improved concordance for Gemini and Claude.

Conclusion

Mainstream consumer large language models showed substantially different performance in postoperative endometrial cancer decision making. Although Gemini achieved higher concordance and fewer major safety issues than ChatGPT and Claude, no model demonstrated performance sufficient to support autonomous clinical use in multidisciplinary management.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.