Skip to content

Author

R. D. Del Hoyo

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Aug 2026

Testing Knowledge Boundaries: Adversarial Evaluation of LLMs for Antimicrobial Stewardship.

OBJECTIVES Evaluate whether general-purpose large language models (LLMs) demonstrate competencies suitable for antimicrobial stewardship (AMS) support and characterize their failure modes. METHODS Cross-sectional evaluation of seven LLMs (GPT-5, Claude Sonnet 4.5, Gemini 2.5 Pro, Grok 4, Llama-3.3-70b-instruct, Qwen 2.5-72b-instruct, DeepSeek-chat-v3.1) using 30 clinical scenarios mapped to ESCMID AMS competency frameworks. Scenarios included deliberate traps for fabrication and dangerous recommendations. Six AMS experts from the Netherlands and Spain performed blinded dual evaluation using content scores (0-5 scale) and binary safety flags for fabrication and danger. Standard and incentivizing prompt framings were compared. RESULTS Four commercial models achieved mean content scores above 3.9/5.0: Claude Sonnet 4.5 (4.06), Gemini 2.5 Pro (3.96), Grok 4 (3.96), and GPT-5 (3.94). Open-weight models scored significantly lower (2.94-3.57). No model achieved more than 63% responses free of fabrication or danger flags. However, fabrication did not impair clinical utility in non-trap scenarios (all within-category comparisons p>0.20). Danger flags ranged from 6.7% to 16.7% across models, with no significant difference between commercial and open-weight models. Incentivizing prompts were associated with a consistent 0.48-point-content score improvement (p=0.006), though significance attenuated after accounting for scenario-level clustering. Evaluators endorsed LLMs as useful AMS support tools with moderate supervision (5/6), identifying documentation preparation and trainee education as promising applications. CONCLUSIONS Medically untrained LLMs demonstrate competencies suitable for supervised AMS support. Fabrication remains the central safety challenge and requires verification workflows; danger, though less frequent (6.7-16.7%), concentrated in identifiable and therefore mitigable failure modes. Non-clinical stewardship tasks (education, documentation, communication) can benefit now, whereas clinical recommendations require expert oversight. Mapping these boundaries allows AMS teams, particularly those understaffed or without on-site infectious diseases expertise, to decide where LLM support adds value rather than risk.

Ángela Abejez-Arrizabalaga, Galadriel Pellejero-Sagastizabal, Rocío Aznar-Gimeno et al. · 0 citations