HIVMedQA is developed, a clinician-curated benchmark of HIV-related open-ended medical question-answer pairs spanning basic knowledge, clinical reasoning, complex patient vignettes, and bias-modified scenarios that provides a structured benchmark for evaluating LLMs in HIV clinical decision support.
Abstract
Large language models (LLMs) are emerging as tools to support clinical decision making. HIV management is a compelling use case due to its complexity and dynamic nature, involving diverse treatment options, comorbidities, and adherence challenges. However, integrating LLMs into clinical practice raises concerns about accuracy, safety, and clinician acceptance. Despite growing interest, their performance in HIV care remains poorly studied, and benchmarking is lacking.
We developed HIVMedQA, a clinician-curated benchmark of HIV-related open-ended medical question-answer pairs spanning basic knowledge, clinical reasoning, complex patient vignettes, and bias-modified scenarios. We evaluated seven general-purpose and three medical LLMs. Performance was assessed using lexical similarity and an extended medical LLM-as-a-judge framework capturing key clinical dimensions, including question comprehension, reasoning, knowledge recall, bias, potential harm, and factual accuracy, to better capture nuances relevant to the medical domain, with additional evaluation by HIV-experienced physicians.
Performance varies substantially across models and task complexity. Gemini 2.5 Pro achieves the highest overall scores, followed by Claude 3.5 Sonnet and MedGemma-27B. Knowledge recall is generally stronger than question comprehension or clinical reasoning. Medical LLMs do not consistently outperform general-purpose models, and model size alone does not predict performance. Several models are sensitive to cognitive bias prompts. LLM-as-a-judge scoring aligns better with clinician assessment than lexical metrics.
HIVMedQA provides a structured benchmark for evaluating LLMs in HIV clinical decision support. Current LLMs show promise, but limitations in reasoning, bias robustness, and safety indicate that careful validation, domain-specific evaluation, and clinician oversight remain essential before clinical deployment.
Large language models (LLMs) have a growing range of applications in medicine, and their evaluation is critical for ensuring they provide benefit and not harm. This evaluation can be more challenging than traditional machine learning for many reasons, including probabilistic and open-ended outputs, and behavior that shifts with prompt design and accumulated context. This review covers four key areas of LLM evaluation: principles of study design, statistical methods, capability evaluation and clinical context evaluation. Capability evaluation considers different benchmarks, including multiple-choice, agentic and multi-turn benchmarks, alongside operational metrics like token usage. Clinical context evaluation addresses establishing accuracy of free text outputs, such as human review and LLM-as-a-judge, and clinical trial approaches. Across sections, we describe underlying concepts and potential pitfalls, while emphasizing the importance of aligning evaluation methods with the research question. Together, this article aims to provide a pragmatic basis for designing and executing rigorous evaluations of healthcare LLMs.
S. E. McKinney, P. Vu, S. Justice et al.· 0 citations
Following clinical decision pathways (CDPs) defined by clinical practice guidelines is essential for safe and reliable medical decision-making. However, existing medical large language model (LLM) benchmarks mainly evaluate final-answer accuracy, providing limited evaluation of models'ability to adhere to guidelines. To address this gap, we introduce MEGA-CDP, a benchmark for evaluating whether medical LLMs can generate guideline-adherent CDPs using provided guidelines as references. MEGA-CDP is constructed from 2,274 English and Chinese clinical practice guidelines through a guideline-to-case pipeline, yielding 42,353 clinical cases with explicit reference CDPs. It supports both single-turn vignette and multi-turn interactive settings, and introduces a CDP-oriented evaluation framework for measuring pathway consistency. Experiments on 16 representative LLMs show that reliable clinical decision support remains challenging for current models, demonstrating the need for CDP-oriented evaluation and the value of MEGA-CDP for advancing guideline adherence in medical LLMs.
Nuo Chen, Xin Jiang, Zi-Long Wang et al.· 0 citations
Introduction: Large language models (LLMs) have the potential to strengthen clinical decision-making in low-resource primary healthcare (PHC) settings. However, most LLMs are developed and benchmarked in high-resource settings and evidence on their safety and contextual appropriateness in Sub-Saharan Africa remains limited. The hypMOOVE-PHC study is the hypothetical vignette phase of the Massive Open Online Validation and Evaluation (MOOVE) initiative, implemented in Kenya, Malawi, and Tanzania. It aims to validate a pool of LLMs through clinical review of expert-generated vignettes. Methods and analysis: This is a fully crossed repeated-measures comparative evaluation study. In each country, experienced clinicians develop 200-250 hypothetical clinical vignettes reflecting realistic patient presentations and independently produce a human benchmark care plan for each. Vignettes are used to prompt a selection of six open-source and proprietary LLMs selected based on code availability, local hostability, and model size. During in-person workshops (valiDATAthons), independent clinical experts rate LLM- and human-generated responses in source-attribution masked side-by-side comparisons across five dimensions (clinical soundness, safety, contextual fit, clarity & completeness, and appropriate confidence). The primary endpoints are each LLM's overall performance profile and non-inferior safety profile, as compared to the human benchmark. At minimum, 358 evaluations per LLM (or 1,253 paired evaluations in total) are required per country. Ethics and dissemination: The study is approved by the EPFL Human Ethics Research Committee in Switzerland, Harvard T.H. Chan School of Public Health in the USA, KNH-UoN Ethics and Research Committee in Kenya, MUBAS Research Ethics Committee in Malawi, and MUHAS Research and Ethics Committee and National Institute for Medical Research in Tanzania. Findings will be reported according to the TRIPOD-LLM framework and shared with national ministries of health, disseminated at conferences and in peer-reviewed journals, and de-identified benchmark data will be released under FAIR principles.
P. Macharia, C. Kachimanga, M. Mahende et al.· medRxiv· 0 citations
A thorough review of the developments in LLM technologies, their uses in clinical and administrative settings, as well as their ethical considerations are reviewed to suggest a conceptual structure for responsible implementation that will ensure both technological innovation and patient safety, as well as regulatory compliance and ethical health care practices.
Noah Wright· International Journal of Mod...· 0 citations
A side effect that misrepresents patients is measured: a one-sentence DEI prompt appended to a medical question leads models to add patient demographic attributes the question never stated, in effect rewriting who the patient is.
Antimicrobial resistance is one of the leading threats to global health. The inappropriate use of antibiotics is one of its strongest drivers. Large language models (LLMs) have entered clinical discussion as tools that might support diagnosis and prescribing. However, their specific role in infectious-disease (ID) care has not been mapped in a structured way. This scoping review charts the breadth, applications, performance signals, and limitations of LLMs used for clinical decision support in ID diagnosis and antimicrobial prescribing. We followed the PRISMA Extension for Scoping Reviews (PRISMA-ScR). We searched PubMed/MEDLINE for peer-reviewed sources that described LLM-based decision support in ID diagnosis or antimicrobial use. Sources were charted by application domain, model evaluated, study design, reported outcomes, and stated limitations. Findings were summarized descriptively. Inferential statistics are reported only as stated by the primary studies. Forty-seven sources were included. They were mapped to five domains: diagnosis and clinical reasoning; antimicrobial prescribing and stewardship; resistance and mechanism prediction; consultation and disease-specific management; and mitigation, evaluation, and ethics. LLMs answered medical-knowledge and case questions at or near passing thresholds. In vignette studies, they achieved diagnostic accuracy comparable to physicians. However, prescribing performance was inconsistent. Agreement with ID specialists on antibiotic choice was often modest, accuracy fell as case complexity rose, and unsafe or guideline-discordant advice recurred. Retrieval-augmented generation and domain grounding consistently improved accuracy and reduced hallucination. Current evidence supports an assistive, human-supervised role for LLMs in ID care rather than autonomous prescribing. Standardized evaluation, prospective validation, local grounding, and explicit stewardship oversight are prerequisites for safe adoption. Not applicable. This study is a scoping review of existing literature and is not a clinical trial; therefore, no clinical trial registration number is applicable. Large language models reason over infectious-disease diagnostic knowledge at a level comparable to physicians in vignette studies. However, their antimicrobial-prescribing performance is inconsistent and degrades as cases become more complex, so they should not prescribe autonomously. Agreement with infectious-disease specialists on antibiotic choice is frequently modest, and models may recommend less-preferred agents or unnecessarily long durations; a clinician must review every recommendation. Correct answers can rest on flawed reasoning, so evaluations must assess transparency and rationale quality, not accuracy alone. Retrieval-augmented generation and grounding in local antibiograms and guidelines consistently improve accuracy and reduce hallucination, and they should be treated as prerequisites for clinical use (Giuffrè et al. 2025; Liu et al. 2025). Safe deployment requires standardized evaluation, prospective validation against real outcomes, and mandatory infectious-disease and antimicrobial-stewardship oversight. To our knowledge, this is the first scoping review to focus specifically on large language models at the intersection of infectious-disease diagnosis and antimicrobial prescribing, rather than on artificial intelligence in medicine broadly. The review charts the evidence across five application domains and contrasts diagnostic performance with prescribing performance. It makes three distinct contributions. First, it separates the comparatively encouraging diagnostic-reasoning literature from the more cautionary prescribing literature, a distinction that is often blurred in general reviews. Second, it positions retrieval-augmented generation and local data grounding as the recurring thread that turns a generic chatbot into a clinically safer decision aid. Third, it synthesizes these findings into a practical, safeguarded framework. That framework places infectious-disease and antimicrobial-stewardship oversight at the center of any deployment, and gives clinicians and informatics staff an explicit map of the opportunities and the unresolved risks.
M. Sannathimmappa· Bulletin of the National Res...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.