Aug 2026· npj Viruses· Vol 4· 0 citations· 31 references
Medicine
TL;DR
A hybrid model in which LLMs generate draft patient information that is subsequently refined by clinical experts is presented, particularly in rapidly evolving therapeutic domains lacking standardized educational resources, to address antimicrobial resistance.
Abstract
Bacteriophage therapy is re-emerging as a potential strategy to address antimicrobial resistance, but standardized patient education materials are limited. Large language models (LLMs) are increasingly used for patient-facing medical information. The quality of LLM-generated responses to 20 patient-relevant questions was evaluated by 12 clinicians and research experts in bacteriophage therapy independently rated each response for accuracy, completeness, clarity, and tone/empathy using 5-point Likert scales. Expert suggestions for improvement were recorded. A total of 960 ratings were analyzed. Adjusted mean scores ranged from 3.36 to 3.96 across domains, indicating generally favorable evaluations for all models. Significant differences among LLMs were observed for completeness and tone/empathy (Holm-adjusted p = 0.042 for both), but not for accuracy or clarity. Differences were small in magnitude (Cohen’s d = 0.12–0.29). Claude scored significantly lower than the other models for completeness and tone/empathy, while Perplexity achieved the highest completeness scores. Experts recommended improvements for 34–40% of responses; wrong information was given in 20%. The best responses were revised into an expert-informed patient guide provided as Supplementary Material, presenting a hybrid model in which LLMs generate draft patient information that is subsequently refined by clinical experts, particularly in rapidly evolving therapeutic domains lacking standardized educational resources.
This scoping review is the first scoping review to focus specifically on large language models at the intersection of infectious-disease diagnosis and antimicrobial prescribing, rather than on artificial intelligence in medicine broadly.
M. Sannathimmappa· Bulletin of the National Res...· 0 citations
LLMs demonstrated similar guideline concordance, suggesting patients can expect comparable accuracy across platforms, and LLMs generally improved FKGL scores compared to the AAO-HNS CPG, they demonstrated lower FRE scores, indicating mixed results on overall readability.
Hetal Lad, Emily S Kwon, Ayushi Chadha et al.· Journal of Otorhinolaryngolo...· 0 citations
On Allergy/Immunology Step 1-style questions, Gemini and Grok demonstrated higher accuracy than ChatGPT, although their overall accuracies remained approximately 81%.
M. Carroll, Sabrina Kentis, Hannah Kareff et al.· Applied Clinical Informatics· 0 citations
HIVMedQA is developed, a clinician-curated benchmark of HIV-related open-ended medical question-answer pairs spanning basic knowledge, clinical reasoning, complex patient vignettes, and bias-modified scenarios that provides a structured benchmark for evaluating LLMs in HIV clinical decision support.
Gonzalo Cardenal-Antolin, J. Fellay, Bashkim Jaha et al.· Communications Medicine· 1 citation
BACKGROUND
Large language models are increasingly investigated as clinical decision support tools, but their reliability for therapeutic drug monitoring interpretation remains poorly explored. Phenytoin and digoxin, two narrow therapeutic index drugs, represent clinically challenging test cases.
OBJECTIVES
To evaluat...
H. Azmakan, Tahmour Azamakan· Journal of the American Phar...· 0 citations
BACKGROUND
large language models (LLMs) are increasingly used for public health information, but their performance in nutrition advice for cirrhosis remains uncertain. We investigated four LLMs in answering public questions about cirrhosis nutrition across safety, accuracy, empathy, information reliability and quality,...
Jun-Zheng Li, Ying-Jie Wu, Man Yang et al.· Nutrición Hospitalaria· 0 citations