Skip to content
Review Open access

Severity Stratification Changes What Hallucination Rates Mean: An Item-Matched Audit of Five Large Language Models in Orthodontic Decision Support

Sep 2026 · Diagnostics · 0 citations · 23 references

Abstract

Background/Objectives: Large language models (LLMs) are consulted for clinical decision support, and their reliability is judged by hallucination prevalence—a measure treating every unsupported element as equivalent, although a fabricated citation and a fabricated protocol differ in what a clinician acting on them would do. We tested whether prevalence tracks clinical risk. Methods: Five LLMs answered a 100-item orthodontic benchmark validated by three external orthodontists (content validity index 0.923). The reference standard was fixed by construction for the 45 items naming a non-existent entity, so any substantive elaboration is unsupported by design; citations were adjudicated against PubMed and CrossRef. Responses were coded with a seven-category taxonomy and assigned to clinical, operational, or epistemic severity tiers in a post hoc exploratory stratification, pre-specified rather than prospectively registered. Because all models answered the same items, comparisons used Cochran’s Q with pairwise McNemar tests and generalised estimating equations clustered on item. Results: Of 500 responses, 449 (89.8%) contained a hallucination but only 65 (13.0%; 95% CI 10.3–16.2) were clinically consequential: prevalence was roughly sevenfold greater than the rate of clinically consequential output as the authors defined it. Between-model differences were large for undifferentiated prevalence (73–100%; Q = 50.51, p < 0.001) and contracted at the clinical tier (10–16%; Q = 10.00, p = 0.040), where no pairwise contrast survived adjustment. Item-level clustering was far stronger for clinically consequential output than for undifferentiated prevalence (intra-class correlation 0.83 versus 0.07). Conclusions: Prevalence and clinically consequential error are not interchangeable and rank models differently. The stratification is exploratory and author-defined, and the benchmark stress-tests susceptibility to fabricated premises rather than surveying natural use.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.