Skip to content
Open access

Assessing safety and trustworthiness of large language models in medicine

Sep 2026 · Communications Medicine · 0 citations

Abstract

The remarkable capabilities of large language models make them increasingly compelling for use in real-world healthcare applications. However, the risks associated with using these artificial intelligence systems in medicine are not systematically understood. The aim of this study is to characterize these risks by applying five key principles for safe and trustworthy medical artificial intelligence: truthfulness, resilience, fairness, robustness, and privacy. We introduce MedGuard-Bench: a safety benchmark featuring one thousand expert-verified questions covering ten specific aspects of our five core principles. We use this comprehensive corpus to systematically evaluate sixteen commonly used large language models, assessing their safety and reliability in medical contexts. We show that current large language models generally perform poorly on most of our safety tests, regardless of their safety alignment mechanisms. Our evaluation demonstrates that these models fall significantly short when compared to the high performance and reliability of human physicians. Despite reports indicating that advanced large language models can match or exceed human performance in various medical tasks, this study reveals a significant safety gap in current technology. This underscores the crucial need for ongoing human oversight and the implementation of strict safety guardrails before deploying these tools in clinical practice.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.