Aug 2026· BMC Cancer· Vol 26· 0 citations· 43 references
Medicine
TL;DR
Both models demonstrated high agreement with expert GIST MTB recommendations, with no significant performance difference between them, and support a potential assistive role for LLMs in GIST MTB workflows, while underscores the continued necessity of expert oversight.
Abstract
Gastrointestinal stromal tumors (GISTs) are molecularly heterogeneous neoplasms whose management depends on individualized, multidisciplinary decision-making. While multidisciplinary tumor boards (MTBs) represent the standard of care, access remains limited in many clinical settings. This study evaluates the performance of two large language models in generating GIST MTB recommendations and assesses their agreement with expert MTB decisions using predefined clinical evaluation criteria. This retrospective single-center study included 99 GIST cases discussed at an institutional MTB. A structured prompt was developed to extract clinical variables and generate treatment recommendations. ChatGPT-5 and Qwen3 were independently evaluated across five predefined domains: diagnostic recommendations, therapeutic modalities, treatment sequence and timing, systemic therapy regimen selection, and clinical contextualization. Two expert reviewers scored all outputs in a blinded fashion. Normalized scores, inter-model comparisons, perfect-case rates, and inter-rater agreement were analyzed. Both models demonstrated high concordance with expert MTB recommendations, with mean total normalized scores of 0.901 for ChatGPT-5 and 0.875 for Qwen3, without a significant difference between models (p > 0.05). Perfect agreement was observed in 52.5% of ChatGPT-5 cases and 48.5% of Qwen3 cases (p > 0.05). Diagnostic recommendations scored significantly lower than all other domains in both models (all adjusted p < 0.05). Overall inter-rater agreement was almost perfect (weighted Cohen’s kappa=0.978). Both models demonstrated high agreement with expert GIST MTB recommendations, with no significant performance difference between them. Diagnostic reasoning represented the weakest domain, reflecting the challenge of reconstructing context-dependent workup decisions from tumor board documentation. These findings support a potential assistive role for LLMs in GIST MTB workflows, while underscoring the continued necessity of expert oversight.
ABSTRACT Objectives Large language models (LLMs) are increasingly proposed as clinical decision‐support tools; however, their agreement with real‐world multidisciplinary tumor board (MDT) decisions remains insufficiently investigated in thyroid oncology. To evaluate the concordance between treatment recommendations gen...
B. B. Büyük, Arzu Or Koca, F. Toprak et al.· Laryngoscope Investigative O...· 0 citations
Abstract Background Hepatopancreatobiliary (HPB) malignancies require complex treatment planning that often relies on multidisciplinary team (MDT) discussions. Large language models (LLMs) have recently been explored for clinical decision support, but their performance within real-world multidisciplinary decision envir...
Jun-Jo Sung, Eui Hyuk Chong, Incheon Kang et al.· Journal of Medical Internet...· 0 citations
Large language models (LLMs) such as GPT-4 are being evaluated for their use as supportive tools in oncological treatment planning. However, in pancreatic cancer, current studies are confined to predefined question–answer formats, while studies specifically investigating real-world scenarios that benchmark LLM performa...
F. Gehrisch, K. Kirkgöz, Antonie Willner et al.· Langenbeck's archives of sur...· 0 citations
ChatGPT achieved the highest overall concordance, although all models generated clinically acceptable recommendations in most cases, and all LLMs demonstrated high concordance with consultant-led thyroid cancer MDT decisions.
A. White, Kerry A. Leyton, N. Patel et al.· Updates in Surgery· 0 citations
Multidisciplinary teams (MDTs) are central to colorectal cancer management, where treatment decisions increasingly depend on the integration of tumor stage, molecular characteristics, patient fitness, and multimodal treatment strategies. However, MDT workflows are time-consuming, subject to inter-team variability, and...
A. Nikitaras, S. M. Tsoti, M. Pramateftakis· Frontiers in Oncology· 0 citations
Historically, precision medicine implies matching one biomarker to cognate monotherapies. However, next-generation precision oncology must address cancer complexity. Indeed, advanced tumors have a median of five genomic alterations; with ~700 cancer-causing genes, there are >1 trillion patterns. We describe an algo...
Ally Perlina, Subha Krishnan, K. Bush et al.· npj Genomic Medicine· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.