Skip to content
Open access

Augmenting Head and Neck Multidisciplinary Tumor Board Recommendations With Locally Run Large Language Models: Prospective Evaluation of Real-World Implementation

Aug 2026 · JMIR AI · Vol 5 · 0 citations · 23 references
Medicine

TL;DR

The data demonstrate that the integration of LLMs in today’s MDT workflow is feasible and may benefit the quality of decision-making in specific cases, and suggests that more advanced local models may offer safe, rapid, and cost-effective support for MDT decision-making.

Abstract

Abstract Background Multidisciplinary tumor boards (MDTs) constitute the foundation of modern tumor therapy. Large language models (LLMs) are widely discussed for optimizing their recommendations. Objective This is the first prospective feasibility study evaluating the implementation of locally run LLMs on real-world cases within a regular head and neck MDT. Methods Seventeen patients participated in the study. The MDT cases were processed by 2 different local LLMs (gemma-3-12b and gpt-oss-20b) to obtain treatment recommendations. The MDT conferred as usual. After the decision was made, the MDT was presented with the LLMs’ recommendations. If deemed to be beneficial, the MDT’s recommendation was adjusted. The MDT members rated the LLMs’ responses inter alia, for medical adequacy on a 6-point Likert scale. In addition, a tabular comparison of the MDT’s and LLMs’ recommendations was carried out. Results In one case, 6% (1/17, 95% CI 0%‐29%), the LLM was able to substantially improve the MDT recommendation by underscoring a follow-up examination that had not yet been performed. Concordance regarding the curative or palliative therapy regimen reached 94% (16/17, 95% CI 71%‐100%); for gemma-3-12b and 59% (10/17, 95% CI 33%‐82%) for gpt-oss-20b. Gemma-3-12b stated the same first-line therapy regimen as the MDT as first-line in 35% (6/17, 95% CI 14%‐62%) of cases, and gpt-oss-20b in 41% (7/17, 95% CI 18%‐67%) of cases. In 59% (10/17, 95% CI 33%‐82%) of patients, gemma-3-12b stated the MDT’s first-line therapy regimen, albeit with a different priority, while for gpt-oss-20b, it was 41% (7/17, 95% CI 18%‐67%) of patients. Medical adequacy, as rated by the MDT members, revealed a median of 5 (IQR 2‐5) for gemma-3-12b and 4 (IQR 3‐5) for gpt-oss-20b. MDT members stated potentially hazardous information in 27% (25/93, 95% CI 18%‐37%) of ratings for gemma-3-12b and 17% (14/83, 95% CI 9%‐26%) of ratings for gpt-oss-20b. Conclusions Locally run LLMs improved the MDT recommendation in 1 case and were primarily useful for identifying potentially relevant missing information in other cases, underscoring that they cannot replace MDTs. However, their observed benefit suggests that more advanced local models may offer safe, rapid, and cost-effective support for MDT decision-making. The study should be seen as an exploratory setting focusing on practical insights rather than benchmarking its clinical impact. Accordingly, the data demonstrate that the integration of LLMs in today’s MDT workflow is feasible and may benefit the quality of decision-making in specific cases.

Read PDF

Similar papers

Open access Sep 2026

Large Language Models in Multidisciplinary Decision-Making for Hepatopancreatobiliary Oncology: Retrospective Comparative Feasibility Study

LLM-generated treatment recommendations demonstrated moderate alignment with MDT decisions in HPB oncology, indicating that concordance alone is insufficient for evaluating LLMs as clinical decision support tools.

Jun-Jo Sung, Eui Hyuk Chong, Incheon Kang et al. · 0 citations
Open access Sep 2026

Concordance between GPT-4 and a multidisciplinary tumor board in pancreatic cancer: A prospective pilot study

GPT-4 demonstrated substantial agreement with MDT recommendations in patients with newly diagnosed or suspected pancreatic cancer, however, specific abstract prompting did not enhance the rate of concordance and GPT-4’s limitations in individualized or complex contexts underscore the need for a cautious future integrat...

F. Gehrisch, K. Kirkgöz, Antonie Willner et al. · 0 citations
Review Open access Sep 2026

Consensus recommendations for the management of locally advanced rectal cancer

Background Locally advanced rectal cancer (LARC) considerably impairs quality of life (QoL) and compromises prognosis. Immunotherapy induces complete responses in most patients with mismatch repair deficient/microsatellite instability-high (dMMR/MSI-h) LARC, enabling organ preservation; however, challenges include acce...

R. Hofheinz, R. García-Carbonero, F. Ciardiello et al. · 0 citations
#large language models Open access Sep 2026

Large Language Models Versus Multidisciplinary Tumor Board Decisions in Thyroid Cancer

While LLMs demonstrate promising concordance in standardized thyroid cancer management, they are best positioned as supportive decision aids—such as in MDT preparation and workflow streamlining—rather than replacements for expert multidisciplinary evaluation, particularly in complex clinical scenarios.

B. B. Büyük, Arzu Or Koca, F. Toprak et al. · 0 citations
Review Open access Aug 2026

Precision oncology meets Generative AI: assessing large language models in multidisciplinary GIST tumor boards

Both models demonstrated high agreement with expert GIST MTB recommendations, with no significant performance difference between them, and support a potential assistive role for LLMs in GIST MTB workflows, while underscores the continued necessity of expert oversight.

Reza Dehdab, Judith Herrmann, Fiona Mankertz et al. · 0 citations
Review Open access Aug 2026

Multidisciplinary Delphi consensus-based recommendations on the use of image-guided radiotherapy for upper gastrointestinal tumours: an AIRO–AITRO–AIFM study

Abstract Objectives Upper gastrointestinal (GI) malignancies represent one of the most challenging scenarios for image-guided radiotherapy (IGRT). Despite technological advances, clinical practice remains heterogeneous among centres. This national, multidisciplinary Delphi consensus, jointly promoted by AIRO, AITRO, an...

Elena Galofaro, P. de Franco, Filippo De Renzi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.