Skip to content
Review Open access

A real‐world analysis of AI chatbot performance for medicines information enquiries

Aug 2026 · Journal of Pharmacy Practice and Research · 0 citations · 24 references

TL;DR

The study results demonstrated the chatbots' responses were typically suboptimal; albeit, a significant minority prompted a follow up to review the chatbot response, and the study conforms with the Declaration of Helsinki.

Abstract

The provision of medicines information (MI) services requires interpretation and clinical judgement of complex scenarios by pharmacists. To date, few studies have assessed the performance of artificial intelligence (AI) chatbots to assist pharmacists providing MI advice. To evaluate the performance and risk associated with two AI chatbots (Microsoft Copilot and Google Gemini) to answer medicines‐related questions. A sample of 20 questions answered by the local MI service in November 2023 was entered in the two chatbot applications in January 2024 (round 1) and May 2024 (round 2). All questions were preceded with the prompt ‘I'm a pharmacist’. Chatbot responses were evaluated by comparing with a reference answer given by the MI service using a consensus process in the domains of content, patient management, risk of patient harm, and follow up review. Ethical approval was granted by the Canterbury District Health Board Research Office (Reference no: 20311) and the study conforms with the Declaration of Helsinki. For the 20 questions answered by both chatbots, few of the round 1 responses ( n = 4 for Copilot and n = 2 for Gemini) were considered complete and with adequate information to commence patient management with no risk of harm. Most were incomplete ( n = 13 for Copilot and n = 15 for Gemini) regarding content, but none were high risk of causing harm. In round 1, four responses from Copilot and eight from Gemini were flagged for follow up review. There was no significant difference in performance between chatbots in round 1 (p = 0.68) or between rounds 1 and 2 (Copilot p = 0.25 and Gemini p > 0.99). Our study results demonstrated the chatbots' responses were typically suboptimal; albeit, a significant minority prompted a follow up to review the chatbot response.

Read PDF

Similar papers

Open access Aug 2026

AI Chatbots as a source of Ramadan medication management advice for patients with diabetes: a multilingual comparative evaluation

While most responses aligned with the IDF-DAR guideline, some harmful recommendations persisted, and response consistency fluctuated over time, the results suggest that AI chatbots should serve as a supplementary resource rather than a substitute for professional medical advice.

S. Alomair, Maryam Alsuwayq, Walaa Alabbad et al. · 0 citations
Jul 2026

Performance of Large Language Model-Based Chatbots in Primary Health Care Teleconsultations: A Comparison Between Human and Artificial Intelligence-Generated Responses.

AI demonstrates substantial potential as a support tool for teleconsultation services despite existing limitations, according to quality criteria, conciseness, coherence, and comprehensibility.

Gabriela Dário Mendes Barros, Carlos Eduardo Menezes Amaral, César Macieira et al. · 0 citations
Open access Aug 2026

Assessing the Accuracy of Artificial Intelligence Chatbots in Medical Information Retrieval: A Structured Query-based Evaluation

Background: Artificial intelligence chatbots are increasingly used to obtain medical and drug-related information, but their accuracy for clinical use remains uncertain. Objective: To evaluate and compare the performance of three large language models—ChatGPT, Gemini, and Grok—in responding to standardised drug-related...

S. Dhohan, Gagan D. Urs, K. Sneha et al. · 0 citations
Review Open access Jul 2026

Development of an AI-driven chatbot for medication-assisted treatment standards in Scotland

This study provides the first prototype of an AI-driven chatbot specifically designed for MAT professionals, demonstrating feasibility of integrating advanced AI technologies to address information access barriers in addiction treatment.

Sandra C. Nwobi, Zainab Loukil, Abbas Jawahar · 0 citations
Review Open access Aug 2026

A Multidisciplinary Evaluation of ChatGPT and DeepSeek's Effectiveness in Answering Questions on Antibiotic Prophylaxis

In recent years, artificial intelligence technologies have become a source of information for patients seeking answers to questions related to medicine and dentistry. Antibiotic prophylaxis is administered prior to invasive dental and medical procedures to prevent systemic complications. Patients who turn to artificial...

Sukran Acipinar, Mehmet Şahinbaş, Irem Aydin et al. · 0 citations
Review Open access Aug 2026

Multilingual Conversational AI Chatbots for Efficient Healthcare Delivery During Case History-Taking: A Systematic Review

Multilingual AI chatbots demonstrate a boost in healthcare efficiency, a reduction in language barriers, and the promotion of health equity, but exhibit challenges regarding validation, workflow integration, and evaluation standards, along with ethical issues such as privacy and bias.

R. Sharanesha, Deepti Virupakshappa, A. Abushanan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.