Skip to content
#small language model Open access

Performance of large language models as a source of clinical information on bacteriophage therapy

Aug 2026 · npj Viruses · Vol 4 · 0 citations · 31 references
Medicine

TL;DR

A hybrid model in which LLMs generate draft patient information that is subsequently refined by clinical experts is presented, particularly in rapidly evolving therapeutic domains lacking standardized educational resources, to address antimicrobial resistance.

Abstract

Bacteriophage therapy is re-emerging as a potential strategy to address antimicrobial resistance, but standardized patient education materials are limited. Large language models (LLMs) are increasingly used for patient-facing medical information. The quality of LLM-generated responses to 20 patient-relevant questions was evaluated by 12 clinicians and research experts in bacteriophage therapy independently rated each response for accuracy, completeness, clarity, and tone/empathy using 5-point Likert scales. Expert suggestions for improvement were recorded. A total of 960 ratings were analyzed. Adjusted mean scores ranged from 3.36 to 3.96 across domains, indicating generally favorable evaluations for all models. Significant differences among LLMs were observed for completeness and tone/empathy (Holm-adjusted p = 0.042 for both), but not for accuracy or clarity. Differences were small in magnitude (Cohen’s d = 0.12–0.29). Claude scored significantly lower than the other models for completeness and tone/empathy, while Perplexity achieved the highest completeness scores. Experts recommended improvements for 34–40% of responses; wrong information was given in 20%. The best responses were revised into an expert-informed patient guide provided as Supplementary Material, presenting a hybrid model in which LLMs generate draft patient information that is subsequently refined by clinical experts, particularly in rapidly evolving therapeutic domains lacking standardized educational resources.

Read PDF

Similar papers

Review Open access Sep 2026

Large language models for clinical decision support in infectious disease diagnosis and antimicrobial prescribing: a scoping review

This scoping review is the first scoping review to focus specifically on large language models at the intersection of infectious-disease diagnosis and antimicrobial prescribing, rather than on artificial intelligence in medicine broadly.

M. Sannathimmappa · 0 citations
Review Open access Aug 2026

Benchmarking Large Language Model Responses Against Surgical Clinical Practice Guidelines for Chronic Rhinosinusitis: The Importance of User Prompts

LLMs demonstrated similar guideline concordance, suggesting patients can expect comparable accuracy across platforms, and LLMs generally improved FKGL scores compared to the AAO-HNS CPG, they demonstrated lower FRE scores, indicating mixed results on overall readability.

Hetal Lad, Emily S Kwon, Ayushi Chadha et al. · 0 citations
Open access Mar 2026

A Comparative Analysis of Large Language Model Performance on USMLE Step 1-Style Allergy/Immunology Questions: Evaluating Correctness and Consistency

On Allergy/Immunology Step 1-style questions, Gemini and Grok demonstrated higher accuracy than ChatGPT, although their overall accuracies remained approximately 81%.

M. Carroll, Sabrina Kentis, Hannah Kareff et al. · 0 citations
Open access Aug 2026

Benchmarking large language models for HIV medical decision support

HIVMedQA is developed, a clinician-curated benchmark of HIV-related open-ended medical question-answer pairs spanning basic knowledge, clinical reasoning, complex patient vignettes, and bias-modified scenarios that provides a structured benchmark for evaluating LLMs in HIV clinical decision support.

Gonzalo Cardenal-Antolin, J. Fellay, Bashkim Jaha et al. · 1 citation
Sep 2026

Performance of large language models on narrow therapeutic index drug monitoring: Implications for clinical pharmacy practice.

BACKGROUND Large language models are increasingly investigated as clinical decision support tools, but their reliability for therapeutic drug monitoring interpretation remains poorly explored. Phenytoin and digoxin, two narrow therapeutic index drugs, represent clinically challenging test cases. OBJECTIVES To evaluat...

H. Azmakan, Tahmour Azamakan · 0 citations
Open access Sep 2026

Performance of large language models in answering public questions about nutrition in cirrhosis: a comparative study.

BACKGROUND large language models (LLMs) are increasingly used for public health information, but their performance in nutrition advice for cirrhosis remains uncertain. We investigated four LLMs in answering public questions about cirrhosis nutrition across safety, accuracy, empathy, information reliability and quality,...

Jun-Zheng Li, Ying-Jie Wu, Man Yang et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.