Skip to content

Performance of large language models on narrow therapeutic index drug monitoring: Implications for clinical pharmacy practice.

Sep 2026 · Journal of the American Pharmacists Association · pp. 103524 · 0 citations
Medicine

Abstract

Background

Large language models are increasingly investigated as clinical decision support tools, but their reliability for therapeutic drug monitoring interpretation remains poorly explored. Phenytoin and digoxin, two narrow therapeutic index drugs, represent clinically challenging test cases.

Objectives

To evaluate three current-generation LLMs on phenytoin and digoxin TDM interpretation across seven clinical reasoning domains, identify domain-specific strengths and limitations, and characterize failure patterns using a qualitative error taxonomy.

Methods

Thirty structured clinical vignettes were submitted to Claude Sonnet 4.6, ChatGPT 5.5, and Gemini 3.1 Pro. The blinded responses were independently scored by two raters using a seven-domain rubric. Between-model differences were assessed using Friedman and post-hoc Wilcoxon signed-rank tests, with Holm adjustment across domain-level tests and Bonferroni correction for pairwise comparisons. All 90 responses underwent error taxonomy analysis. Second independent responses were subsequently generated and scored using the same procedure to assess agreement across repeated generations.

Results

Claude achieved the highest mean performance (94.9% ± 8.3%), significantly outperforming ChatGPT (84.5% ± 12.9%, p<0.001) and Gemini (83.4% ± 11.9%, p=0.001). All models showed high performance on level interpretation and toxicity assessment, but significant between-model differences emerged on pharmacokinetic reasoning, management, monitoring, and uncertainty acknowledgment (all Holm-adjusted p≤0.02). Overconfident reasoning was the most common error category (45.4% of total error appearances). Repeat-generation analysis showed high test-retest agreement across generations (ICC=0.90; 95% CI, 0.82-0.94).

Conclusion

The evaluated LLMs performed strongly in recognizing TDM problems in structured vignettes but showed variability in reasoning-intensive tasks, with overconfident reasoning representing a potential patient safety concern. Pharmacist oversight remains essential for TDM tasks.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.