Large Language Models and Retrieval-Augmented Platforms for the Diagnosis and Management of Periodontal Diseases: A Blinded Expert-Rated Comparative Study of 11 Systems.
Abstract
Aim
To compare retrieval-augmented systems with general-purpose large language models (LLMs) on standardised periodontal clinical vignettes.
Materials And Methods
Eleven AI systems were evaluated: nine general-purpose LLMs, one general-purpose retrieval-augmented platform (Perplexity) and one medical-domain retrieval-augmented platform (OpenEvidence). Each responded to 30 synthetic vignettes covering acute, chronic and complex periodontal scenarios. Six blinded periodontists scored responses on a 5-point Likert scale for accuracy, safety, freedom from hallucinations and completeness in a randomised block design. Friedman and Conover-Iman tests with Holm correction were applied; mixed-effects and ordinal models served as sensitivity analyses.
Results
At least one parameter scored dangerous (≤ 2) in 3.3%-46.7% of responses across platforms, despite mean composite scores (3.28-4.86) exceeding the rubric midpoint of 3.0. Between-model differences were significant (p < 0.001), with a small-to-medium overall effect (Kendall's W = 0.17) and large within-category effects (W up to 0.82). Perplexity, OpenEvidence and Claude 4.7 Opus formed a top tier.
Conclusion
Retrieval-augmented systems rated highest, but this advantage was confounded with response length. The dangerous-response spread argues against undifferentiated use. These tools should assist, not replace, specialist judgement.