PURPOSE
Large language models (LLMs) have proven performance for certain diagnostic tasks; however, limited studies have evaluated their consistency in recommending appropriate medication regimens for a given diagnosis. Medication management is a complex task that requires synthesis of drug formulation and complete order instructions for safe use. Here, the performance of GPT-4o, an LLM available with OpenAI's ChatGPT, was tested on 3 medication management tasks.
METHODS
GPT-4o performance was tested on 3 medication tasks: identifying available formulations for a given generic drug name, identifying drug-drug interactions (DDIs) for a given medication regimen, and preparing a medication order for a given generic drug name. For each experiment, the model's raw text response was captured exactly as returned and evaluated using clinician evaluation in addition to standard LLM metrics, including Term Frequency-Inverse Document Frequency (TF-IDF) vectors, normalized Levenshtein similarity, and Recall-Oriented Understudy for Gisting Evaluation (ROUGE-1/ROUGE-L) F1 score between each response and its reference string.
RESULTS
For the first task of drug-formulation matching, GPT-4o had 49% accuracy for generic medications being matched to all available formulations, with an average of 1.23 omissions per medication and 1.14 hallucinations per medication. For the second task of drug-drug interaction identification, the accuracy was 54.7% for identifying the DDI pair. For the third task, GPT-4o generated order sentences containing no medication or abbreviation errors in 65.8% of the cases.
CONCLUSION
Model performance for basic medication tasks was consistently poor. This evaluation highlights the need for domain-specific training through clinician-annotated datasets and a comprehensive evaluation framework for benchmarking performance.
Kelli Henry, Steven Xu, Kaitlin Blotske et al.· American Journal of Health-S...· 0 citations
BACKGROUND
Accurate detection of drug-drug interactions (DDIs) is a fundamental component of safe medication management. Traditional rule-based clinical decision support systems for DDI identification lack higher-order reasoning and contribute to alert fatigue. Large language models (LLMs) have potential for DDI identification but may hallucinate, inconsistently identify interactions, and provide overly confident responses despite uncertainty. Prior studies have emphasized accuracy, but few have examined whether LLM uncertainty expression aligns with error risk.
OBJECTIVE
To evaluate LLM performance in DDI identification using a clinician-validated dataset and to assess whether prompt-based mitigation strategies improve knowledge-aware uncertainty expression, defined as alignment between refusal behavior and likelihood of error.
METHODS
We developed a clinician-curated DDI identification task consisting of 250 medication lists, each containing 1 clinically relevant interacting drug pair, to evaluate 5 LLMs: GPT-5-Chat, GPT-4o-mini, Gemma-27B, LLaMA3-70B, and Qwen3-32B. Models were evaluated using 3 prompt formats and a zero-shot approach, with no task-specific training or examples provided. Prompts were designed to encourage uncertainty acknowledgment, including a patient safety-focused mitigation prompt to support clinically appropriate and cautious responses. Each case was run 9 times per prompt condition. The primary outcome was the Refusal Index (RI), which quantifies alignment between model refusal behavior and likelihood of error. Secondary outcomes included overall accuracy, accuracy given attempted, refusal rate, self-consistency, F score, weighted score, and entropy.
RESULTS
Across models, overall DDI identification accuracy ranged from 54.1% to 83.7%. GPT-5-Chat demonstrated the highest overall accuracy and self-consistency, whereas Quen3-32B demonstrated the lowest accuracy but the highest refusal rates. Alignment between refusal behavior and likelihood of error was weak to moderate (RI range 0.104-0.574) and varied by model. Prompt-based mitigation strategies produced inconsistent effects on RI and did not reliably recalibrate uncertainty behavior. Notably, higher overall accuracy and response stability did not consistently correspond to stronger knowledge-aware uncertainty expression. Qwen3-32B and GPT-4o-mini increased refusal rates under mitigation prompting but not in situations where responses were more likely to be incorrect.
CONCLUSIONS
Substantial variability exists in both DDI identification performance and uncertainty calibration across LLMs. Prompt design alone was insufficient to consistently improve knowledge-aware uncertainty expression. Because safe deployment of LLMs in medication management depends not only on accuracy but also on appropriate deferral when error risk is elevated, multidimensional evaluation frameworks are essential before clinical use.
A. Tilley, Brian Murray, K. Henry et al.· Journal of Managed Care & Sp...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.