Comparative Evaluation of Large Language Models in Answering Patient Questions Following Periodontal and Peri-Implant Examination: An Expert-Based Study.
Aug 2026· Journal of Stomatology Oral and Maxillofacial Surgery· pp.
102956
· 0 citations· 29 references
Medicine
TL;DR
G Gemini demonstrated superior clinical precision and safety, whereas Claude provided more comprehensive and readable explanations, which support the integration of LLMs as pragmatic, high-ecological-validity complementary tools for patient education, while emphasizing the persistent necessity for professional clinical oversight.
Abstract
INTRODUCTION
Large language models (LLMs) are increasingly utilized for medical and dental information retrieval, yet their ability to interpret authentic, patient-style inquiries remains insufficiently investigated. This study compared the performance of ChatGPT, Claude, and Gemini in responding to patient-oriented queries related to periodontal and peri-implant diseases.
Materials And Methods
Unlike traditional investigations using expert-generated questions, this study employed 40 realistic, patient-oriented queries designed to simulate the post-examination cognitive state, blending colloquial language with partially retained clinical jargon. Each query was submitted to GPT-4o, Claude Sonnet 5 and Gemini 2.5 Pro generating 120 responses. Three blinded periodontists independently evaluated scientific accuracy, completeness, clinical safety, and overall quality using a 5-point Likert scale. Automated text analysis assessed readability metrics (Flesch Reading Ease, Flesch-Kincaid Grade Level, Gunning Fog Index) and linguistic characteristics. Statistical protocols included Friedman, Bonferroni-adjusted Wilcoxon signed-rank, and Model Dominance analyses.
Results
Significant performance differences were observed among the models across all expert-rated domains (all p < 0.001). Gemini achieved the highest expert ratings for scientific accuracy (4.77 ± 0.22), clinical safety (4.94 ± 0.15), and overall quality (4.85 ± 0.20), and was identified as the most frequently top-ranked platform via dominance analysis. Conversely, Claude performed significantly better regarding response completeness (4.77 ± 0.22) and demonstrated the most favorable overall readability profile, yielding the lowest Flesch-Kincaid Grade Level (7.54 ± 1.27). GPT-4o consistently received the lowest expert ratings across all evaluated domains.
Discussion
While all evaluated LLMs generated high-quality responses to realistic periodontal queries, their functional strengths were highly multidimensional. Gemini demonstrated superior clinical precision and safety, whereas Claude provided more comprehensive and readable explanations. These findings support the integration of LLMs as pragmatic, high-ecological-validity complementary tools for patient education, while emphasizing the persistent necessity for professional clinical oversight.
While LLMs show strong potential in supporting dental education through standardised exams, their performance varies by model and question type, and further improvements are needed to enhance reliability across different dental disciplines.
N. Acar, Fatih Sengul, Periş Çelikel et al.· European journal of dental e...· 0 citations
BACKGROUND
Artificial intelligence-powered large language models (LLMs) are increasingly used by patients seeking quick information regarding dental and medical problems. Despite their growing popularity, concerns remain regarding the accuracy, clarity, and clinical usefulness of LLM-generated responses. This study aim...
Sajad Jahantigh, R. Amid, A. Moscowchi et al.· Clinical Advances in Periodo...· 0 citations
The expert-guided visual correction framework introduced here provides a generalizable, mechanism-based approach to benchmarking multimodal AI diagnostic performance that extends beyond periodontics to other visually driven diagnostic domains in medicine.
There may be significant differences in how effectively LLMs support patients with surgical queries, particularly in areas needing detailed explanation, and usually required minimal clarification in areas needing detailed explanation.
T. Davis, B. Guevel, K. Logishetty et al.· Annals of the Royal College...· 0 citations
It is demonstrated that LLMs can serve as auxiliary tools in providing orthodontic patient information but responses should be checked by an expert before being presented to patients and should be adapted into simpler and more understandable language.
Alperen Erdoğan, Orhan Çiçek· Journal of International Den...· 0 citations
In this single-generation benchmark, the sampled outputs showed different profiles in specialist-rated quality, clinical content, potential clinical risk, readability, and externally observable response latency under the specific interface configurations tested.
I. Karaca, Esmanur Başer, Emre Ulubaş et al.· Journal of Clinical Medicine· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.