Skip to content
Conference Open access

Factual Hallucination in Medical Large Language Models: Typology, Evaluation, and Systematic Governance

Sep 2026 · Exploring Science Academic Conference Series · 0 citations · 26 references

Abstract

Large language models (LLMs) have demonstrated transformative potential in clinical documentation generation, diagnostic assistance, and patient consultation. However, their tendency toward “hallucination”— generating semantically fluent but factually inconsistent content with established medical knowledge or input context—constitutes a core safety barrier to clinical deployment. This paper systematically reviews the typology, evaluation frameworks, underlying mechanisms, and mitigation strategies for medica l LLM hallucinations. First, we propose a fine-grained hallucination classification framework based on medical knowledge graphs and establish a three-tier risk stratification scheme that references FDA medical AI software risk classification standards. Sec ond, we analyze multi-factor causes from data, model, and inference dimensions, highlighting the unique characteristics of the medical domain. Third, we systematically review current evaluation benchmarks and methodologies, deeply analyzing the limitations of automated metrics and the potential of LLMs as evaluators. Regarding mitigation strategies, we propose a three-stage classification framework—”training-stage internal intervention —inference-stage external constraints —post-processing collaborative verification” —that provides in-depth analysis of each strategy's internal mechanisms and trade- offs. Finally, we discuss challenges in real-world clinical deployment, emphasizing that the “human-in-the- loop” paradigm remains an irreplaceable safety line.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.