Jul 2026· Philosophical transactions. Series A, Mathematical, physical, and engineering sciences· Vol 384 2324· 1 citation· 42 references
Medicine
TL;DR
Five widely used prompting strategies across five influential LLMs in the latest medical bias benchmark reveal substantial heterogeneity in both effectiveness and overhead across models, with no strategy proving universally effective and some even exacerbating bias.
Abstract
Large language models (LLMs) demonstrate expert-level performance in various medical scenarios, yet their outputs can exhibit bias against groups or individuals with specific sensitive attributes, posing risks to patient safety and undermining trust in LLMs for healthcare. Recent research suggests that prompt engineering offers a convenient way to adjust model outputs, with the potential to mitigate such biases. However, there is a lack of empirical studies that systematically examine the effectiveness of prompt engineering and its trade-offs among fairness, accuracy and inference overhead. To fill this gap, we empirically evaluate five widely used prompting strategies across five influential LLMs in the latest medical bias benchmark. Results reveal substantial heterogeneity in both effectiveness and overhead across models, with no strategy proving universally effective and some even exacerbating bias. Chain-of-thought prompting yields the largest reduction, lowering the average gap across all scenarios by 2.4 percentage points, where the largest reduction is 6.2 percentage points, obtained in the DeepSeek-V3.1-sex case. Furthermore, the results of the McNemar test also show that it achieves the largest number of significant bias reduction cases (8/15), primarily by improving performance on unprivileged groups. These findings provide practical guidance for the fair deployment of LLMs in healthcare and highlight that mitigating medical bias remains a challenging problem requiring sustained efforts from both the artificial intelligence (AI) and medical communities. To support future research on fair AI in healthcare, we shall release all results and source code. This article is part of the theme issue 'Safe, secure and robust AI for safety-critical systems'.
A side effect that misrepresents patients is measured: a one-sentence DEI prompt appended to a medical question leads models to add patient demographic attributes the question never stated, in effect rewriting who the patient is.
Medical large language models are often judged by how many clinical questions they answer correctly. That view is useful, but it misses a practical risk. A model may know the right answer and still change its response when the same case is written in a different patient voice. This paper evaluates that risk as SDoH aware narrative anchoring bias. We use NarrativeShield SDoH MedQA, a counterfactual medical question answering dataset in which each case appears in persona based narratives while the answer key remains fixed. The dataset is reshaped from wide format into case grouped persona rows. We evaluate three open source instruction tuned LLMs from the Qwen2.5 family: 1.5B, 3B, and 7B. The final experiment uses 300 clinical cases and produces 8,100 model responses across three prompting conditions. We report persona level accuracy, counterfactual consistency, correct consistency, and narrative sensitivity error. Qwen2.5 7B achieves the best accuracy at 56.33 percent and the best correct consistency at 40.33 percent. Paired McNemar exact tests show significant accuracy gains for 7B over 3B in all prompt settings. Even so, narrative sensitivity remains, with the lowest error still at 31.67 percent. These results suggest that trustworthy clinical decision support should be evaluated by both average correctness and stability across medically equivalent patient narratives.
Background: LLM judges increasingly score whether clinical language models give overconfident answers under incomplete evidence, yet whether a measured"safety gain"reflects real behavior change or the judge's calibration is unresolved. Using a structured evidence-sufficiency prompt as a test case, we asked whether it reduces unsafe overconfident answers, how far that effect depends on the scoring judge, and what it costs in helpfulness. Methods: In a retrospective public-data benchmark (Real-POCQi, HealthBench, MedRBench), four models (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Flash, Grok 4.3) answered a fully paired common panel (1,200 cells) with a standard prompt and the wrapper. The pre-specified endpoint was the paired reduction in unsafe overconfidence scored by the primary judge (GPT-5.4-nano); secondary analyses added a different-family judge (Claude Sonnet 5), a correctness judge, matched scaffold controls, and a blinded three-clinician review. Results: Unsafe overconfidence fell from 49.3% to 24.7%, a paired reduction of 24.7 points (95% CI 21.8-27.7; p<0.001), robust in direction across models and paraphrases. Magnitude was judge-dependent: Sonnet agreed on direction but nearly halved the effect (+13.1 points), with one-directional disagreement. Blinded clinicians characterized the primary judge as a high-sensitivity (1.00), low-specificity (0.55) screen, not a calibrated rate. The gain carried a model-specific helpfulness cost (correct diagnosis 80.3% to 50.3%): near-free for GPT-5.5, near-total for Gemini (-58 points). Matched scaffold controls showed genuine behavior change, not judge circularity. Conclusions: LLM-judged clinical safety effects should be reported as directional and relative, anchored to human review and evaluated jointly with helpfulness, not as calibrated absolute rates. This does not establish clinical deployment readiness.
Retrieval-augmented large language models (LLMs) promise source-linked clinical support, but their value depends on whether displayed evidence guides rather than distorts physician reliance. We developed CORA, an agentic retrieval-augmented LLM, to investigate how source-linked assistance affects physician decision-making. CORA maintained benchmark performance and achieved larger gains on cases published after the models'training-data cutoffs. In a study of 46 physicians, accuracy increased from 70.8% unaided to 82.6% with CORA. Supporting citations predicted correct answers (87.7% vs 65.5%), but citations created an important asymmetry: perceived support increased adoption of correct advice from 34% to 76.9% but when an incorrect LLM answer appeared citation-supported, physician resistance to it fell from 92% to 34.8%. These findings show that source-linked LLM assistance can improve physician accuracy while introducing a grounding-dependent safety risk.
T. Chanda, C. Wies, Franziska Schramm et al.· 0 citations
Abstract
Medication errors remain among the most significant preventable causes of patient harm worldwide, contributing to increased morbidity, mortality, prolonged hospitalization, and escalating healthcare expenditure. Large Language Models (LLMs) — including GPT-4, Gemini, Claude, and Llama — have emerged as promising clinical decision-support tools capable of interpreting complex medical terminology, analyzing prescriptions in real time, and flagging potential errors before medications reach the patient. This review synthesizes current evidence on the application of LLMs in prescription error detection, presents a consolidated system architecture and operational workflow for LLM-enabled medication safety pipelines, and critically examines their benefits, limitations, and future trajectory. Evidence from recent clinical evaluations indicates that LLM-based decision-support tools can achieve high concordance with expert pharmacist judgment and measurably reduce near-miss medication events when deployed with appropriate safeguards. However, challenges including AI hallucination, data privacy, algorithmic bias, regulatory ambiguity, and the continued necessity of human oversight must be addressed before widespread clinical adoption. The review concludes that LLMs hold substantial promise as complementary — rather than autonomous — decision-support systems capable of transforming medication safety and pharmacy practice.
Keywords: Large Language Models; Prescription Error Detection; Medication Safety; Clinical Decision Support; Artificial Intelligence in Healthcare; Electronic Health Records; Pharmacovigilance
K. T. K. Kumar, Koyya Gowtham Reddy, K. Reddy· International Scientific Jou...· 0 citations
It is found that LLM access enhances performance on standardized clinical vignettes in all three countries, and policymakers should prioritize structured integration of LLMs as decision-support tools, combined with targeted training, local validation, and safeguards against automation bias rather than relying on access alone.
N. Rounding, L. S. Arif, Janine Berg et al.· 0 citations