On complex ID scenarios, large language models responses were variable and caution is required when deploying these models in ID domains without specialist oversight, suggesting caution is required when deploying these models in ID domains without specialist oversight.
Abstract
Objectives: Large language models (LLMs) are increasingly used in medicine, but evaluation is often on multiple choice questions and management of common conditions. Infectious diseases (ID) can present complex scenarios that require considerations beyond guideline-based responses. We assessed LLM performance in these situations including with ID-specific criteria to consider infection control or antimicrobial stewardship (AMS). Methods: We evaluated four LLMs (Claude 3.5 Sonnet, GPT-4o, GPT-o1, and a local instance of Llama 3.1 8B) in October 2024, on five complex ID vignettes. The LLM responses were each evaluated for 18 items by two board-certified ID clinicians and pairwise comparisons were performed between LLMs. Results: There was no significant difference between performance of GPT-o1, GPT-4o and Claude Sonnet on general medical criteria, and were comparable with respect to how often they provided an unsafe response (GPT-o1 30%, GPT-4o 40%, Claude 37%) and contained a critical omission (GPT-o1 27%, GPT-4o 43%, Claude 47%). Llama 3.1 8B had significantly decreased performance for most criteria. On ID-specific criteria, GPT-o1 outperformed other models and all models significantly outperformed Llama for interpreting microbiology results, AMS principles, appropriate antimicrobial spectrum and infection control considerations. Performance was poor in secondary prevention and management of risk factors. Conclusions: On complex ID scenarios, LLM responses were variable. The open-source, smaller Llama 3.1 8B model performed poorly and large, non-reasoning models varied, but more than 30% of responses containing a risk of harm or critical omission. These findings suggest caution is required when deploying these models in ID domains without specialist oversight.
Background: Fever in the returning traveller is a common but challenging presentation with a broad, geography-dependent differential diagnosis. Timely assessment can be difficult for front-line clinicians, especially outside tropical medicine settings. Large language models (LLMs) may support clinical decision-making by generating differential diagnoses, but comparative evaluation workflows remain underdeveloped.
Objective: To develop and evaluate an LLM workflow to assess the reliability and efficacy of LLM-generated differential diagnoses for written fever-in-returning-traveller cases, compare performance across four LLMs, and secondarily evaluate LLM case-generation capability.
Methods: We studied 21 travel-related diagnoses using four LLMs (ChatGPT 5.2 Thinking, GPT-o3, Llama 3.2, and Mistral Instruct). Infectious diseases fellows created cases across the diagnoses (n=84; 4/diagnosis), which were used to calibrate AI case-generation prompting. Each LLM then generated matched cases (n=84; 21/model). Clinician and AI-generated cases were pooled, assigned unique IDs, and randomized so each analysis model received equal numbers of clinician and AI cases (n=168 analyses; 42/model). Feasibility outcomes were completion, retries/system errors, and logged response time. Expert qualitative scoring is underway.
Results: We completed dataset generation and model-output acquisition for 252 outputs across 21 diagnoses. Case-generation retries occurred in 17/84 outputs (20.2%), all first-pass refusals were from Llama 3.2, all resolved with repeat prompting. Case-analysis retries were rare (1/168, 0.6%; single no-response, resolved), with no persistent failures. Balanced allocation was achieved. Median case-analysis response times for cloud-based models were 11s (ChatGPT 5.2 Thinking) and 4s (GPT-o3).
Conclusions: Preliminary findings demonstrate a feasible, scalable, and reproducible workflow for comparative LLM evaluation in fever in the returning traveller assessment. Model-specific refusal behaviour is an important implementation consideration. Ongoing expert qualitative scoring will determine comparative reliability and efficacy. Current conclusions are limited to operational feasibility and dataset generation reliability.
Bhavya Gandhi, Leo Morjaria, Lori Israelian et al.· McMaster University Medical...· 0 citations
Background: Antimicrobial resistance poses a major threat to global public health. Large language models (LLMs) offer new possibilities for optimizing antibiotic prescribing decisions, but the capabilities of general-purpose versus domain-specific medical LLMs under different prompting strategies remain to be clarified. Methods: This double-blind, randomized-sequence evaluation used a 2X2 factorial design comparing four AI conditions-the domain-specific model MedGo and the general-purpose model DeepSeek V3.5, each under standard direct prompting and chain-of-thought (CoT) prompting-alongside real physician prescriptions across 59 complex inpatient infection cases. Five parallel regimens were generated per case and independently evaluated by three senior clinicians (1-5 comprehensive score and five domain sub-scores). ChatGPT 5.2 was additionally assessed as an automated evaluation tool. Results: Score ranking: real physicians > MedGo-CoT > DeepSeek-CoT > MedGo> DeepSeek (Friedman test, p<0.001). In base mode, MedGo significantly outperformed DeepSeek (Holm-adjusted p=0.040). CoT improved both models (Holm-adjusted p<0.001 for DeepSeek; p=0.024 for MedGo) and reduced score dispersion. MedGo-CoT significantly outperformed DeepSeek-CoT in individualized adjustment (adjusted p<0.001) and dosing precision (adjusted p=0.005). ChatGPT-expert correlation was negligible (overall Kendall {tau}=0.153, p=0.003; subgroup {tau}=0.06-0.20, all p>0.05). Conclusions: Domain-specific medical LLMs enhanced by CoT approach the antibiotic decision-making level of real physicians, with advantages in individualization and dosing precision. However, notable deficiencies persist in antimicrobial stewardship ecological awareness and automated evaluation reliability, underscoring the continued indispensability of senior clinical expertise.
Y. Liu, C. Zhang, F. Wang et al.· medRxiv· 0 citations
BACKGROUND
Early decision-making in acute pancreatitis (AP) involves diagnostic confirmation, early severity triage, escalation thresholds, and initiation of guideline-concordant management under time pressure and incomplete information. Large language models (LLMs) may support structured bedside reasoning, but their clinical usefulness cannot be inferred from guideline knowledge alone.
METHODS
A cross-sectional, scenario-based comparative evaluation was conducted in January 2026 using 20 AP scenarios: 15 refined hypothetical vignettes and 5 de-identified, privacy-modified real-life case patterns. GPT-4, GPT-5, and Gemini received identical single-turn prompts. Model access was through OpenAI API gpt-4-0613, OpenAI API gpt-5, and Google Vertex AI Gemini 1.0 Pro; temperature was set to 0.0, and each prompt was repeated three times per model. Outputs were scored by two independent clinician-raters using a prespecified 1-5 ordinal rubric across guideline concordance, safety, actionability, and data-synthesis quality. Two senior board-certified surgeons independently generated expert reference pathways for comparison.
RESULTS
GPT-5 achieved the highest guideline concordance (4.28 ± 0.38) and safety (4.20 ± 0.45) profiles. GPT-4 provided the clearest stepwise actionability (4.15 ± 0.48), whereas Gemini showed the strongest data-synthesis quality (4.22 ± 0.52). With deterministic settings, internal consistency across three repeated runs was 100%. All models demonstrated clinically relevant failure modes, particularly unwarranted certainty under missing data; this occurred in 12/20 GPT-4, 7/20 GPT-5, and 15/20 Gemini outputs.
CONCLUSION
No model should be used as a stand-alone bedside decision-maker for AP. In this scenario-based early evaluation, GPT-5 was the most safety-aligned model, GPT-4 was the most operationally actionable, and Gemini was strongest for synthesis, but all require clinician oversight, prospective validation, and governance before clinical deployment.
Y. K. Çalışkan, Fatih Başak, Olgun Erdem· World Journal of Surgery· 0 citations
Anaplastic thyroid cancer (ATC) is a rare, aggressive malignancy with poor prognosis. Adherence to guidelines from the National Comprehensive Cancer Network (NCCN), American Thyroid Association (ATA), and European Society for Medical Oncology (ESMO) is critical for optimal patient outcomes. As large language models (LLMs) increasingly enter clinical workflows, rigorous evaluation of their alignment with established guidelines is essential. We evaluated five leading LLMs for their ability to generate guideline-concordant responses to clinical questions about ATC. We conducted a comparative study in 2025 following TRIPOD-LLM (Transparent Reporting of a Multivariable Model for Individual Prognosis or Diagnosis, Large Language Models) guidelines. Seventy clinical questions of varying complexity were developed from ATA, NCCN, and ESMO guidelines. Three surgical oncology experts validated each question and subsequently evaluated responses from five LLMs: ChatGPT 4.1, ChatGPT 5, Gemini 2.5 Pro, Claude Sonnet 4, and DeepSeek R1. Each response was scored for relevancy, clarity, accuracy, and adequacy on a 5-point Likert scale. Inter-rater reliability was assessed using both intraclass correlation coefficients (ICC) and Gwet’s AC2 with ordinal weights. Model comparisons used the Kruskal-Wallis test with Dunn’s post-hoc analysis and Bonferroni correction. A pre-specified sensitivity analysis excluding the unblinded model (ChatGPT 5) was performed to confirm robustness. Significant performance differences emerged across all four metrics: accuracy (p = 0.007), adequacy (p = 0.003), clarity (p = 0.014), and relevance (p < 0.001). Gemini 2.5 Pro achieved the highest median accuracy (4.5), followed by DeepSeek R1 (4.4), while ChatGPT 4.1 scored lowest (4.0). ICC values ranged from 0.34 to 0.44 (poor to moderate), but Gwet’s AC2 yielded substantially higher estimates of 0.61 to 0.73 (moderate to substantial agreement), reflecting the impact of restricted score range on conventional reliability metrics. The sensitivity analysis excluding ChatGPT 5 confirmed the performance hierarchy among blinded models, with significance preserved or strengthened across all four metrics. Leading LLMs show variable capacity to align with ATC clinical guidelines. While top-performing models hold promise as supportive tools, their inconsistencies across domains and complexity levels preclude autonomous clinical use. These models should serve strictly as decision aids under expert supervision.
Mohamed Yasser, Ghada Barakat, S. Awny et al.· Scientific Reports· 0 citations
Abstract Background Large language models (LLMs) are increasingly applied in clinical decision support, yet their diagnostic performance in Chinese-language settings and under realistic clinical workflows remains unclear. In particular, how LLMs perform across diseases with different prevalence and under stepwise diagnostic processes has not been well characterized. Objective This study aimed to evaluate the diagnostic capabilities of LLMs for common diseases and rare diseases using clinical vignettes within a hypothetico-deductive framework and to identify their potential and limitations for clinical diagnosis. Methods We evaluated 4 Chinese LLMs (Doubao 1.5, DeepSeek-V3, Kimi K1.5, and Leftdoctor GPT 3.5) using 56 clinical cases (28 chronic obstructive pulmonary disease [COPD], and 28 relapsing polychondritis [RP]) sourced from the China Clinical Case Results Database (March 31-April 14, 2025). Patient information was provided incrementally, starting with the initial medical history, followed by physical examination, and laboratory results. Evaluation metrics included top-3 accuracy (RTop3D), top-1 accuracy (RTopD), final diagnostic accuracy (RFA), and mean reciprocal rank (MRR). Statistical analysis was performed using generalized estimating equations (GEE), Friedman tests, and Wilcoxon signed-rank tests with Bonferroni correction. In addition, a qualitative analysis was conducted to characterize recurrent patterns of diagnostic errors. Results LLMs demonstrated significantly higher diagnostic accuracy for COPD compared to RP across all metrics (P<.001). Diagnostic accuracy improved after additional clinical information was provided, with the improvement mainly observed in RP cases. In RP, diagnostic accuracy increased from 32.14% to 71.43% for DeepSeek and from 35.71% to 78.57% for Doubao, whereas COPD accuracy remained consistently high across all diagnostic stages (82.14%‐92.86%). For COPD, ranking performance was high and comparable among all models (MRR range: 0.82‐0.89; P=.71). In RP, diagnostic performance differed significantly among models (MRR range: 0.10‐0.39; P<.001). Qualitative analysis showed that COPD errors were mainly related to a failure to recognize specific features, whereas RP errors involved more diverse patterns, particularly the neglect of negative evidence and the failure to recognize specific features. Conclusions Chinese LLMs demonstrated relatively strong diagnostic performance for common diseases such as COPD, but lower and less stable performance for rare diseases such as RP. Additional clinical information improved diagnostic accuracy primarily in RP cases, although differences between models remained evident under diagnostically complex conditions. Error patterns in RP cases suggest that current LLMs remain limited in their ability to integrate complex clinical information and exclusionary findings. Careful evaluation and appropriate clinical oversight remain important for their application in clinical practice.
Jiayi Wang, Jiao Yang, Rui Guo· Journal of Medical Internet...· 0 citations
OBJECTIVES
Evaluate whether general-purpose large language models (LLMs) demonstrate competencies suitable for antimicrobial stewardship (AMS) support and characterize their failure modes.
METHODS
Cross-sectional evaluation of seven LLMs (GPT-5, Claude Sonnet 4.5, Gemini 2.5 Pro, Grok 4, Llama-3.3-70b-instruct, Qwen 2.5-72b-instruct, DeepSeek-chat-v3.1) using 30 clinical scenarios mapped to ESCMID AMS competency frameworks. Scenarios included deliberate traps for fabrication and dangerous recommendations. Six AMS experts from the Netherlands and Spain performed blinded dual evaluation using content scores (0-5 scale) and binary safety flags for fabrication and danger. Standard and incentivizing prompt framings were compared.
RESULTS
Four commercial models achieved mean content scores above 3.9/5.0: Claude Sonnet 4.5 (4.06), Gemini 2.5 Pro (3.96), Grok 4 (3.96), and GPT-5 (3.94). Open-weight models scored significantly lower (2.94-3.57). No model achieved more than 63% responses free of fabrication or danger flags. However, fabrication did not impair clinical utility in non-trap scenarios (all within-category comparisons p>0.20). Danger flags ranged from 6.7% to 16.7% across models, with no significant difference between commercial and open-weight models. Incentivizing prompts were associated with a consistent 0.48-point-content score improvement (p=0.006), though significance attenuated after accounting for scenario-level clustering. Evaluators endorsed LLMs as useful AMS support tools with moderate supervision (5/6), identifying documentation preparation and trainee education as promising applications.
CONCLUSIONS
Medically untrained LLMs demonstrate competencies suitable for supervised AMS support. Fabrication remains the central safety challenge and requires verification workflows; danger, though less frequent (6.7-16.7%), concentrated in identifiable and therefore mitigable failure modes. Non-clinical stewardship tasks (education, documentation, communication) can benefit now, whereas clinical recommendations require expert oversight. Mapping these boundaries allows AMS teams, particularly those understaffed or without on-site infectious diseases expertise, to decide where LLM support adds value rather than risk.
Ángela Abejez-Arrizabalaga, Galadriel Pellejero-Sagastizabal, Rocío Aznar-Gimeno et al.· Clinical Microbiology and In...· 0 citations