Large language models for clinical decision support in infectious disease diagnosis and antimicrobial prescribing: a scoping review
Antimicrobial resistance is one of the leading threats to global health. The inappropriate use of antibiotics is one of its strongest drivers. Large language models (LLMs) have entered clinical discussion as tools that might support diagnosis and prescribing. However, their specific role in infectious-disease (ID) care has not been mapped in a structured way. This scoping review charts the breadth, applications, performance signals, and limitations of LLMs used for clinical decision support in ID diagnosis and antimicrobial prescribing. We followed the PRISMA Extension for Scoping Reviews (PRISMA-ScR). We searched PubMed/MEDLINE for peer-reviewed sources that described LLM-based decision support in ID diagnosis or antimicrobial use. Sources were charted by application domain, model evaluated, study design, reported outcomes, and stated limitations. Findings were summarized descriptively. Inferential statistics are reported only as stated by the primary studies. Forty-seven sources were included. They were mapped to five domains: diagnosis and clinical reasoning; antimicrobial prescribing and stewardship; resistance and mechanism prediction; consultation and disease-specific management; and mitigation, evaluation, and ethics. LLMs answered medical-knowledge and case questions at or near passing thresholds. In vignette studies, they achieved diagnostic accuracy comparable to physicians. However, prescribing performance was inconsistent. Agreement with ID specialists on antibiotic choice was often modest, accuracy fell as case complexity rose, and unsafe or guideline-discordant advice recurred. Retrieval-augmented generation and domain grounding consistently improved accuracy and reduced hallucination. Current evidence supports an assistive, human-supervised role for LLMs in ID care rather than autonomous prescribing. Standardized evaluation, prospective validation, local grounding, and explicit stewardship oversight are prerequisites for safe adoption. Not applicable. This study is a scoping review of existing literature and is not a clinical trial; therefore, no clinical trial registration number is applicable. Large language models reason over infectious-disease diagnostic knowledge at a level comparable to physicians in vignette studies. However, their antimicrobial-prescribing performance is inconsistent and degrades as cases become more complex, so they should not prescribe autonomously. Agreement with infectious-disease specialists on antibiotic choice is frequently modest, and models may recommend less-preferred agents or unnecessarily long durations; a clinician must review every recommendation. Correct answers can rest on flawed reasoning, so evaluations must assess transparency and rationale quality, not accuracy alone. Retrieval-augmented generation and grounding in local antibiograms and guidelines consistently improve accuracy and reduce hallucination, and they should be treated as prerequisites for clinical use (Giuffrè et al. 2025; Liu et al. 2025). Safe deployment requires standardized evaluation, prospective validation against real outcomes, and mandatory infectious-disease and antimicrobial-stewardship oversight. To our knowledge, this is the first scoping review to focus specifically on large language models at the intersection of infectious-disease diagnosis and antimicrobial prescribing, rather than on artificial intelligence in medicine broadly. The review charts the evidence across five application domains and contrasts diagnostic performance with prescribing performance. It makes three distinct contributions. First, it separates the comparatively encouraging diagnostic-reasoning literature from the more cautionary prescribing literature, a distinction that is often blurred in general reviews. Second, it positions retrieval-augmented generation and local data grounding as the recurring thread that turns a generic chatbot into a clinically safer decision aid. Third, it synthesizes these findings into a practical, safeguarded framework. That framework places infectious-disease and antimicrobial-stewardship oversight at the center of any deployment, and gives clinicians and informatics staff an explicit map of the opportunities and the unresolved risks.