Stakeholder role-prompting fundamentally alters clinical decisions and ethical value frameworks of frontier LLMs, with the insurer role producing systematic denial of physician-endorsed, patient-preferred treatments.
Abstract
Background: Large language models (LLMs) are increasingly deployed in healthcare, where they may adopt different stakeholder perspectives, yet the effect of role-prompting on clinical ethical reasoning remains poorly characterized. Methods: We evaluated three frontier LLMs: Claude Opus 4.6, GPT-5.4, and Gemini 3.1 Pro across 25 ethically complex medical cases. Each model responded from three stakeholder perspectives (physician, patient, insurer) across three independent runs (675 total responses). Decisions were benchmarked against a six-physician panel. Ethical value prioritization was analyzed from physician- and LLM-provided ranked values. A Patient-Centric Decision Index (PCDI) was developed to quantify LLM decision alignment with patient-preferred out-comes. Results: Among 20 cases with clear physician consensus, LLMs prompted as an insurer reduced alignment with physician majority by 50% for GPT-5.4 (p = 0.004), 45% for Gemini 3.1 Pro (p = 0.008) and 10.5% (NS) for Opus 4.6. The insurer role shifted primary ethical values from beneficence (27%) to financial stewardship (20%) across all LLMs. Conclusions: Stakeholder role-prompting fundamentally alters clinical decisions and ethical value frameworks of frontier LLMs, with the insurer role producing systematic denial of physician-endorsed, patient-preferred treatments. These findings raise the need for standardized LLM patient-centricity benchmarks, and physician oversight when LLMs are deployed in clinical decision-making.
It is found that LLM access enhances performance on standardized clinical vignettes in all three countries, and policymakers should prioritize structured integration of LLMs as decision-support tools, combined with targeted training, local validation, and safeguards against automation bias rather than relying on access alone.
N. Rounding, L. S. Arif, Janine Berg et al.· 0 citations
Large Language Models (LLMs) achieve strong results on many medical benchmarks, but their clinical reasoning remains difficult to evaluate reliably. A central risk is an evaluation illusion: fluent and well-structured explanations can appear clinically convincing even when the final diagnosis is incorrect. We introduce CLExEval, a human-in-the-loop framework for evaluating LLM clinical reasoning under progressive information masking. CLExEval combines 5,600 expert-physician annotations with 200 clinical reasoning traces derived from 40 rare diagnostic cases. Our analysis identifies three recurring failure patterns: (i) verbosity bias, where GPT-4o-mini's diagnostic accuracy drops from 95.0% to 32.5% under information scarcity; (ii) a hidden knowledge paradox, where a specialist model reaches 92.5% maximum diagnostic potential but fails to retrieve that knowledge reliably in verbose contexts; and (iii) a 68.6% reasoning-to-output mismatch, where correct diagnoses appear in reasoning traces but are not reflected in final answers. We further evaluate the LLM-as-a-Judge paradigm on a human-verified failure set (n = 142). GPT-4o-mini approved 47.9% of clinically incorrect outputs, while HuatuoGPT-o1 approved all validly scored failures and showed a positive self-preference bias. These results suggest that standalone automated clinical evaluations can substantially overestimate clinical reliability without expert-grounded validation.
Abin Roy, Afthab Salam Kanniyan, Jawadh Abdul Kabeer et al.· arXiv.org· 0 citations
Multiple-choice medical benchmarks are increasingly saturated, and recent rubric-based evaluations such as HealthBench have shown that open-ended clinical performance is far from solved - its"Hard"subset top score remains 32%. We present a small, deliberately difficult evaluation dataset of five clinician-authored clinical scenarios spanning four specialties (anaesthesia, internal/family medicine, emergency medicine, and obstetrics), each accompanied by an atomic, weighted, MECE rubric (25-62 criteria per task; 184 criteria total) authored from a clinician-drafted golden answer. We evaluate three frontier models: GPT 5.4, Claude Opus 4.7, and Gemini 3.1 Pro. Mean rubric pass rates were 0.47 (Claude), 0.39 (GPT), and 0.37 (Gemini). The central finding is an inversion of clinical priority: the highest-weighted (weight-5, critical) criteria passed at only 32.4-41.7%, while low-stakes weight-1 criteria passed at 80-90%. 56 of 108 critical (weight-5) criteria (52%) were satisfied by no model. Three LLM autoraters reproduced expert met/not-met labels on 92.8-94.7% of 552 graded criteria. We position this as a methods-and-preliminary-findings contribution: the five tasks demonstrate a scalable, defensible pipeline ready to develop into a large-scale benchmark.
Samiha A. Ismail, Fan X. Chen, Ali Merali· 0 citations
LLMs hold substantial potential to enhance healthcare teamwork by supporting clinical decisions, streamlining administrative workflows, and improving patient communication, however, ethical, legal, and accountability concerns remain.
Ilse Super, Olya Rezaeian, Onur Asan· International Journal of Med...· 0 citations
A dual-view approach that connects clinical practice with computational methods is presented, establishing a five-level competency scheme following Miller’s Pyramid and linking deductive, inductive, and abductive reasoning patterns to common medical goals and tasks.
GPT-5 reliably adapts linguistic style to clinical personas but produces limited specialty-specific output diversity, supporting its role as a decision-support adjunct rather than an autonomous specialist simulator.
Derna Stifini, A. Della Penna, André L. Mihaljevic et al.· npj Digital Medicine· 0 citations