Skip to content
Preprint

Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning

Aug 2026 · 0 citations · 27 references
Computer Science

TL;DR

MedPIC-Bench makes conditional rule application measurable and highlights the limitations of static medication-safety accuracy for assessing patient-specific reliability among medical-specific LLMs, whose average CF performance trails that of general LLMs.

Abstract

Applying a valid medication-safety rule when its patient-specific conditions are not met can produce an incorrect decision. Existing medical evaluations largely use isolated and fixed scenarios. A model may therefore answer correctly by recalling a drug-risk association without showing that it used patient information to decide whether the rule applies. To address this gap, we introduce MedPIC-Bench, a benchmark of source-verifiable recommendations and expert-validated questions for patient-specific medication-safety reasoning. It combines guideline-following questions with paired counterfactual questions in which a controlled change in patient information changes whether a rule applies. The benchmark contains 467 questions annotated along six clinical and reasoning dimensions. Across 28 medical-specific, general, and proprietary LLMs, every model performs worse on counterfactual questions, with mean accuracy falling from 63.6\% to 45.1\%. Models perform well when an explicit patient attribute directly signals a familiar contraindication, but struggle when patient information must narrow or withdraw a safety warning. Model rationales often acknowledge the changed patient information, yet the final answers retain the previous safety judgment. This vulnerability persists among medical-specific LLMs, whose average CF performance trails that of general LLMs. MedPIC-Bench therefore makes conditional rule application measurable and highlights the limitations of static medication-safety accuracy for assessing patient-specific reliability.

View source

Similar papers

Review Aug 2026

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR

CliniCARE-Bench is the first deployment-oriented clinical-agent benchmark to jointly evaluate real longitudinal EHR investigation, claim-level evidence grounding, governing-policy use, process adherence, and calibrated abstention within a common patient-level adjudication framework.

Veronica Chatrath, Bryan Zhu, George Pu et al. · 1 citation
Review Open access Aug 2026

Counterfactual Analysis of Executable Clinical Decision Logic

Clinical recommendations are often expressed in narrative form, which limits their direct execution, auditability, and patient-specific interpretation. This paper presents a hybrid decision-support framework that combines Decision Model and Notation (DMN), survey-weighted rule-ensemble learning, and counterfactual sens...

C. Maleki, Y. Bertrand, F. Gailly · 0 citations
Book Open access Aug 2026

RxGuard: Knowledge-Guided Safety Guardrails for Medication Recommendation

Medication combination recommendation predicts a set of medications for a patient visit from longitudinal electronic health records (EHRs) and is safety-critical, since feasibility depends on set-level constraints such as drug--drug interactions (DDIs). Despite substantial progress, existing approaches suffer from thre...

Hang Wang, Hang Dong, Lu Liu · 0 citations
#artificial intelligence Review Sep 2026

Untangling the Mechanisms of Misleading Context in Medical Question Answering

Large language models now answer medical questions with expert-level performance. However, the context these systems act on can be misleading, and misleading context can corrupt a model's medical judgment. To understand how misleading context corrupts this judgment, we examine the model's susceptibility to the context,...

R. Linzmayer, N. Elhadad · 0 citations
Review Open access Sep 2026

Uncertain Pretest Probabilities in Diagnostic Reasoning: The Prevalence Threshold as a Tipping Point.

RATIONALE Post-test probabilities are reported as precise numbers even though the pretest probabilities they depend on are uncertain and, for an individual patient, unobservable. Clinicians need a way to tell the mathematical question-how strongly is a change in the pretest estimate carried through to the post-test pro...

Maya Nadler, J. Balayla · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.