Assessing the Accuracy of Artificial Intelligence Chatbots in Medical Information Retrieval: A Structured Query-based Evaluation
Abstract
Background: Artificial intelligence chatbots are increasingly used to obtain medical and drug-related information, but their accuracy for clinical use remains uncertain. Objective: To evaluate and compare the performance of three large language models—ChatGPT, Gemini, and Grok—in responding to standardised drug-related queries concerning three commonly prescribed drugs. Methods: Three commonly prescribed drugs—metformin, hydrochlorothiazide, and azithromycin—were selected for assessment. Each model was asked ten standardised questions per drug (two questions in each of five categories: indications, off-label indications, drug–drug interactions, adverse drug events, and drug availability). Responses were manually assessed against standard clinical references, principally UpToDate, and scored on a four-point scale from 0 to 3, where 3 represented a completely accurate and clinically sound response. A non-perfect score (0–2) was considered an error. Each model answered 30 questions in total. Results: ChatGPT and Gemini each produced 18 perfect responses, corresponding to an empirical probability of 0.60 for a completely correct answer. Grok produced 14 perfect responses, corresponding to an empirical probability of 0.47. Error rates varied across drugs and models, ranging from 30% to 60%. A two-way analysis of variance (ANOVA) of mean error rates showed that drug type had a statistically significant effect (F = 7.75, p = 0.0421), whereas the effect of the AI model was not statistically significant at the conventional threshold (F = 4.00, p = 0.1111). Grok's numerically higher mean error rate (53.33%, compared with 40.00% for ChatGPT and Gemini) was consistent with its lower empirical success probability. Conclusion: These findings indicate that freely available AI chatbots may provide rapid drug information but show variable accuracy. As a practical implication for current use, rather than a proposed future research direction, their responses should be verified against authoritative clinical references before use in healthcare education or practice.