AI Hallucination And Fabricated References: A Growing Crisis For Medical Researchers A Narrative Review
Abstract
Large language models (LLMs) have introduced a new research integrity threat into the biomedical literature i.e. fabricated references that appear authentic but correspond to no existing publication. This narrative review distinguishes fabrication, the invention of an entirely non-existent source, from unfaithfulness, the citation of a genuine source in support of a claim it does not contain, and argues that the two failure modes require different detection strategies. Early evaluations of ChatGPT-3.5 reported fabrication rates as high as 69% for medical questions, and although newer models perform better, residual rates in the range of 15 to 20% remain in current systems. The review examines five interconnected domains: the autoregressive architecture that makes fabrication structurally likely rather than incidental; the marketing of AI-powered search tools, which frames speed and accuracy as complementary rather than competing, and which disproportionately misleads early-career researchers in resource-limited settings; the compounding effect of AI paraphrasing tools, which can sever the link between a claim and its supporting citation even when no new reference is fabricated; the cascading consequences that extend from manuscript rejection to contamination of the clinical evidence base; and to the current, fragmented state of detection and editorial policy. A recurring finding across the reviewed literature is that professional presentation does not reliably track factual accuracy, so reviewer confidence in AI-assisted text is a poor substitute for verification. The review identifies real-time reference verification integrated into AI-assisted writing workflows as the central unmet need, and proposes a three-tier response spanning researcher-level verification habits, editorial audit infrastructure, and discipline-specific AI literacy training. The findings support the conclusion that fabricated references in medical research reflect a structural property of current language models rather than a transient limitation, and that addressing the problem requires coordinated action rather than reliance on model improvement alone.