Skip to content
Open access

Role prompting modulates linguistic style but not clinical decision structure in GPT-5 tumour board simulation

Aug 2026 · npj Digital Medicine · Vol 9 · 0 citations · 48 references
Medicine

TL;DR

GPT-5 reliably adapts linguistic style to clinical personas but produces limited specialty-specific output diversity, supporting its role as a decision-support adjunct rather than an autonomous specialist simulator.

Abstract

Multidisciplinary tumour boards (MDTs) are the standard for gastrointestinal oncological decision-making but remain resource-intensive. Whether specialty-specific role prompting induces genuinely distinct clinical reasoning in large language models (LLMs)—or merely role-appropriate language around an invariant output—has not been systematically tested. We applied five zero-shot prompting frameworks and a majority-vote ensemble to GPT-5 across 100 gastrointestinal oncology cases with MDT-validated decisions: a simulated MDT, multi-expert deliberation, three specialist personas, and a majority-vote ensemble. Concordance with MDT recommendations ranged from 78% to 87%, with no significant inter-framework differences (Cochran’s Q = 8.46, p = 0.133). Specialty-characteristic language was near-universal (97–100%) but uncorrelated with accuracy. Embedding analysis revealed high semantic similarity across personas (cosine similarity 0.805–0.836; η² = 0.049), contrasting with substantially greater output separation under multi-expert deliberation (η² = 0.554–0.581). GPT-5 reliably adapts linguistic style to clinical personas but produces limited specialty-specific output diversity, supporting its role as a decision-support adjunct rather than an autonomous specialist simulator.

Read PDF

Similar papers

Open access Jul 2026

A multidimensional benchmarking framework for large language models in oncologic decision making

Large language models (LLMs) are increasingly explored as clinical decision support tools in oncology; however, reliance on isolated metrics has limited the development of multi-dimensional evaluation frameworks. This comparative observational study utilized five stepwise, clinically realistic non-small cell lung cancer scenarios reflecting real-world diagnostic, therapeutic, and follow-up decision-making. Open-ended clinical questions were answered by three LLMs (Gemini 2.5 Pro, GPT-5, and Claude Opus 4.1) via their official APIs and compared with evidence-based reference answers. Model outputs were evaluated using expert-rated clinical accuracy and explainability, alongside operational metrics including cost, response time, and generative efficiency. All dimensions were integrated into an expert-weighted Composite Performance Score (CPS). Across 30 clinical questions, significant inter-model differences were observed for all metrics (p < 0.001). GPT-5 achieved the highest accuracy, explainability, and generative efficiency, while Gemini 2.5 Pro demonstrated the lowest cost and Opus 4.1 the fastest response times. Integrated analysis yielded the highest CPS for GPT-5, followed by Gemini 2.5 Pro and Opus 4.1 (Kendall’s W = 0.87). A multi-dimensional evaluation framework integrating clinical quality and operational efficiency provides more actionable insights than single metric assessments, enabling pragmatic model selection for oncology practice. Nevertheless, the use of LLMs in this domain should remain clinician-supervised.

M. Halıcı, Serkan Saltürk, Irem Sayin et al. · 0 citations
Open access Aug 2026

Decoding high-order clinical correlations: a knowledge-driven large language model framework for specialized medical decision-making

Specialized thoracic-surgery questions require the integration of multi-factor clinical relationships within text, yet general-purpose large language models (LLMs) may underperform on such exam-style benchmarks. We constructed a DK-LLM agent by embedding curated medical textbook knowledge into a LangChain-based framework to support domain-specific reasoning. The model was tested on a 56-item thoracic-surgery examination question set in a restricted text-only setting without internet browsing or external tools and was compared with generic LLM baselines, three thoracic surgeons, and three non-expert engineers. Examination score and error patterns were assessed. The knowledge-augmented DK-LLM configuration showed an 11.8-point examination-score advantage over the version without the local knowledge base. Commercial LLM-based agents outperformed the open-source baselines and non-expert participants on this question set, whereas experienced thoracic surgeons achieved the highest scores overall; in the ablation analysis, removing the local knowledge base reduced the examination score by 11.8 percentage points. Embedding domain-specific knowledge into LLMs may improve performance on specialized exam-style thoracic-surgery questions on this text-only benchmark. However, the present 56-item evaluation does not establish clinical equivalence, diagnostic accuracy in practice, multimodal competence, or readiness for real-world clinical decision support.

Qian Li, Yongxin Li, Chao Ye et al. · 0 citations
Open access Aug 2026

Concordance Between Clinical Practice Recommendations Generated by Generative Artificial Intelligence and the Vía RICA 2026 Enhanced Recovery Guideline: A Proof-of-Concept Study Using a Closed Evidence Corpus

Clinical practice guidelines require expert synthesis that large language models (LLMs) might partly automate, yet their ability to reproduce clinically actionable recommendations is poorly quantified. We evaluate an LLM (Claude Sonnet 4.6) against the 103 recommendations of the Spanish enhanced-recovery guideline Vía RICA 2026, grouped in 17 bundles. The model used the panel’s own closed corpus (617 documents) in a multilingual retrievalaugmented generation pipeline. Concordance was assessed twice: by optimal 1:1 bipartite matching (Hungarian) on cosine similarity, and by an LLM-as-a-judge clinical adjudicator (Claude Haiku 4.5) validated against a three-clinician panel (Fleiss’ κ = 0.538). The two schemes bracket a micro F1 of 0.61–0.69 and reveal four findings: (i) a systematic granularity bias, producing 1–8 recommendations per bundle regardless of ground-truth size; (ii) failure of cosine similarity to discriminate within narrow clinical domains; (iii) high reference-concordance precision (0.70–0.81) despite low exhaustiveness; and (iv) no transfer of the GRADE fields, evidence level agreeing no better than chance and strength systematically downgraded. An eight-fold larger retrieval budget left it intact. A corpus audit found 25 documents that formulate recommendations; excluding them lowers judged micro F1 to 0.602. The results delimit the current utility of generative AI for guideline development.

Andrea Moral, Antonio Arroyo, Juan Aparicio et al. · 0 citations
Review Open access Aug 2026

Explainability of decoder-only clinical large language models: A scoping review.

Findings show that clinical LLM explainability has shifted toward fluent generative rationales, but evidence that such explanations reflect model reasoning remains limited, and three regulatory priorities are highlighted: prioritizing explanations that enable independent verification or logic auditing over plausibility-only rationales; preferring inspectable models where regulatory documentation is required; and prospectively validating explanations in clinical workflows before scaling.

Nishant Mishra, Ameen Abu-Hanna, Iacer Calixto · 0 citations
Open access Aug 2026

DSAI-12 ACCURACY, SAFETY, AND READABILITY OF PUBLIC-FACING LARGE LANGUAGE MODELS IN CNS METASTASIS

Abstract Public-facing large language models (LLMs) are increasingly used by patients to obtain medical information, yet their accuracy, safety, and readability in the context of central nervous system (CNS) metastases remain poorly characterized. Fifteen simulated patient questions regarding CNS brain metastases were submitted to four LLMs (ChatGPT, Claude, Gemini, and Open Evidence). Responses were evaluated using a 5-point Likert scale for accuracy based on National Comprehensive Cancer Network (NCCN) guidelines. Incorrect responses were defined as scores ≤2. Readability was assessed using Flesch Reading Ease (FRE), Flesch–Kincaid Grade Level (FKGL), and Gunning Fog index. Word count was also analyzed. Mixed-effects models were used to compare performance across models, accounting for repeated measures by question. ChatGPT demonstrated the highest mean accuracy score (4.8), followed by Open Evidence (4.7), Gemini (4.1), and Claude (3.7). Incorrect responses occurred most frequently with Claude (26.7%), followed by Gemini (13.3%), while no incorrect responses were observed for ChatGPT or Open Evidence. In ordinal regression analysis, ChatGPT and Open Evidence demonstrated superior performance compared to Claude and Gemini (p < 0.01), with no significant difference between ChatGPT and Open Evidence. Readability analysis revealed that Gemini produced the most readable responses (FRE 43.8; FKGL 12.3; Gunning Fog 15.0), followed by ChatGPT, while Claude and Open Evidence generated significantly less readable outputs (p < 0.01). Gemini also generated the longest responses (mean 467 words), whereas ChatGPT and Open Evidence produced shorter responses (∼370 words). LLM performance varied substantially across accuracy, safety, and readability. ChatGPT and Open Evidence achieved the highest accuracy with no incorrect responses, whereas other models, despite greater readability, were more likely to generate incorrect information. Although contemporary LLMs increasingly reflect medical consensus for CNS metastases, inconsistent reliability remains a concern, underscoring the need for caution in patient use.

Michael Fiorino, Mei Hainline, Tanay Poddar et al. · 0 citations