MedPIC-Bench makes conditional rule application measurable and highlights the limitations of static medication-safety accuracy for assessing patient-specific reliability among medical-specific LLMs, whose average CF performance trails that of general LLMs.
Abstract
Applying a valid medication-safety rule when its patient-specific conditions are not met can produce an incorrect decision. Existing medical evaluations largely use isolated and fixed scenarios. A model may therefore answer correctly by recalling a drug-risk association without showing that it used patient information to decide whether the rule applies. To address this gap, we introduce MedPIC-Bench, a benchmark of source-verifiable recommendations and expert-validated questions for patient-specific medication-safety reasoning. It combines guideline-following questions with paired counterfactual questions in which a controlled change in patient information changes whether a rule applies. The benchmark contains 467 questions annotated along six clinical and reasoning dimensions. Across 28 medical-specific, general, and proprietary LLMs, every model performs worse on counterfactual questions, with mean accuracy falling from 63.6\% to 45.1\%. Models perform well when an explicit patient attribute directly signals a familiar contraindication, but struggle when patient information must narrow or withdraw a safety warning. Model rationales often acknowledge the changed patient information, yet the final answers retain the previous safety judgment. This vulnerability persists among medical-specific LLMs, whose average CF performance trails that of general LLMs. MedPIC-Bench therefore makes conditional rule application measurable and highlights the limitations of static medication-safety accuracy for assessing patient-specific reliability.
CliniCARE-Bench is the first deployment-oriented clinical-agent benchmark to jointly evaluate real longitudinal EHR investigation, claim-level evidence grounding, governing-policy use, process adherence, and calibrated abstention within a common patient-level adjudication framework.
Veronica Chatrath, Bryan Zhu, George Pu et al.· 1 citation
Clinical recommendations are often expressed in narrative form, which limits their direct execution, auditability, and patient-specific interpretation. This paper presents a hybrid decision-support framework that combines Decision Model and Notation (DMN), survey-weighted rule-ensemble learning, and counterfactual sens...
C. Maleki, Y. Bertrand, F. Gailly· medRxiv· 0 citations
Medication combination recommendation predicts a set of medications for a patient visit from longitudinal electronic health records (EHRs) and is safety-critical, since feasibility depends on set-level constraints such as drug--drug interactions (DDIs). Despite substantial progress, existing approaches suffer from thre...
Hang Wang, Hang Dong, Lu Liu· Proceedings of the 32nd ACM...· 0 citations
It is suggested that trustworthy clinical decision support should be evaluated by both average correctness and stability across medically equivalent patient narratives.
Large language models now answer medical questions with expert-level performance. However, the context these systems act on can be misleading, and misleading context can corrupt a model's medical judgment. To understand how misleading context corrupts this judgment, we examine the model's susceptibility to the context,...
RATIONALE
Post-test probabilities are reported as precise numbers even though the pretest probabilities they depend on are uncertain and, for an individual patient, unobservable. Clinicians need a way to tell the mathematical question-how strongly is a change in the pretest estimate carried through to the post-test pro...
Maya Nadler, J. Balayla· Journal of Evaluation In Cli...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.