Confident but unsupported: auditing large language models against supplied evidence boundaries in drug-induced liver injury assessment
Abstract
Introduction Large language models (LLMs) are entering drug-safety and regulatory workflows, yet their behavior at the boundary between unavailable evidence and risk reassurance remains poorly characterized. Drug-induced liver injury (DILI) is a stringent setting because evidence is fragmented across labels, case reports, mechanistic studies, and curated knowledge bases, while unsupported low-risk reassurance can be consequential. Methods We evaluated five LLMs on a fixed 47-drug DILI risk-assessment panel using 1,410 parsed responses from closed-book answering and an evidence-gated protocol that restricted responses to supplied PubMed-derived evidence. Curated DILI resources were excluded from prompts and used only for evaluation. After filtering, 38 drugs had no direct DILI decision-support evidence in the supplied packet, and 9 had direct DILI-relevant evidence; a post hoc PubMed title/abstract recall stress test identified five additional drugs with recoverable direct DILI evidence outside the packet. Results In the no-direct-evidence slice defined by the supplied packet, closed-book models rarely abstained, with drug-level abstention ranging from 5.3% to 33.3%; the evidence-gated protocol required abstention, which all models followed for every no-direct-evidence drug. The same pattern held for recent or low-recognition drugs, where evidence-gated abstention reached 92.0% to 100.0% vs. 8.0% to 49.3% under closed-book answering. Closed-book models also produced high-confidence low-risk responses for DILI-positive drugs, a label-discordant pattern largely removed by evidence gating. Independent expert review of selected responses showed that label discordance did not always imply a clinically unreasonable low-risk category, but identified unsafe reassurance through overconfident wording and under-cautious responses in selected cases. When direct DILI evidence was provided, all models preserved citation-grounded non-abstaining answers. However, they differed in how often they committed to a conclusive rather than an uncertain risk category. Citation-bearing evidence-gated responses cited only the supplied PubMed identifiers and achieved 91.2% to 100.0% concordance with the supplied grade. Discussion These findings identify unsupported reassurance as measurable evidence-boundary behavior in LLM drug-risk assessment and establish a reproducible framework for auditing adherence to an externally supplied evidence boundary, defined by PubMed evidence classification and enforced by prompt policy rather than inferred independently by the model.