Skip to content

Comparative Evaluation of Large Language Models in Answering Patient Questions Following Periodontal and Peri-Implant Examination: An Expert-Based Study.

Aug 2026 · Journal of Stomatology Oral and Maxillofacial Surgery · pp. 102956 · 0 citations · 29 references
Medicine

TL;DR

G Gemini demonstrated superior clinical precision and safety, whereas Claude provided more comprehensive and readable explanations, which support the integration of LLMs as pragmatic, high-ecological-validity complementary tools for patient education, while emphasizing the persistent necessity for professional clinical oversight.

Abstract

INTRODUCTION Large language models (LLMs) are increasingly utilized for medical and dental information retrieval, yet their ability to interpret authentic, patient-style inquiries remains insufficiently investigated. This study compared the performance of ChatGPT, Claude, and Gemini in responding to patient-oriented queries related to periodontal and peri-implant diseases.

Materials And Methods

Unlike traditional investigations using expert-generated questions, this study employed 40 realistic, patient-oriented queries designed to simulate the post-examination cognitive state, blending colloquial language with partially retained clinical jargon. Each query was submitted to GPT-4o, Claude Sonnet 5 and Gemini 2.5 Pro generating 120 responses. Three blinded periodontists independently evaluated scientific accuracy, completeness, clinical safety, and overall quality using a 5-point Likert scale. Automated text analysis assessed readability metrics (Flesch Reading Ease, Flesch-Kincaid Grade Level, Gunning Fog Index) and linguistic characteristics. Statistical protocols included Friedman, Bonferroni-adjusted Wilcoxon signed-rank, and Model Dominance analyses.

Results

Significant performance differences were observed among the models across all expert-rated domains (all p < 0.001). Gemini achieved the highest expert ratings for scientific accuracy (4.77 ± 0.22), clinical safety (4.94 ± 0.15), and overall quality (4.85 ± 0.20), and was identified as the most frequently top-ranked platform via dominance analysis. Conversely, Claude performed significantly better regarding response completeness (4.77 ± 0.22) and demonstrated the most favorable overall readability profile, yielding the lowest Flesch-Kincaid Grade Level (7.54 ± 1.27). GPT-4o consistently received the lowest expert ratings across all evaluated domains.

Discussion

While all evaluated LLMs generated high-quality responses to realistic periodontal queries, their functional strengths were highly multidimensional. Gemini demonstrated superior clinical precision and safety, whereas Claude provided more comprehensive and readable explanations. These findings support the integration of LLMs as pragmatic, high-ecological-validity complementary tools for patient education, while emphasizing the persistent necessity for professional clinical oversight.

View source

Similar papers

Open access Sep 2026

Evaluating the Accuracy of Large Language Models in Dentistry: A Multi-Model Study Using Clinical Questions From Turkey's Dental Specialty Exams.

While LLMs show strong potential in supporting dental education through standardised exams, their performance varies by model and question type, and further improvements are needed to enhance reliability across different dental disciplines.

N. Acar, Fatih Sengul, Periş Çelikel et al. · 0 citations
Review Open access Sep 2026

Assessing the accuracy and usability of artificial intelligence-based language models in responding to common periodontal patient questions.

BACKGROUND Artificial intelligence-powered large language models (LLMs) are increasingly used by patients seeking quick information regarding dental and medical problems. Despite their growing popularity, concerns remain regarding the accuracy, clarity, and clinical usefulness of LLM-generated responses. This study aim...

Sajad Jahantigh, R. Amid, A. Moscowchi et al. · 0 citations
Open access Aug 2026

Expert-Guided Visual Correction for Characterizing Diagnostic Performance and Error Patterns of Multimodal Large Language Models Using Periodontal In-Service Examination Images

The expert-guided visual correction framework introduced here provides a generalizable, mechanism-based approach to benchmarking multimodal AI diagnostic performance that extends beyond periodontics to other visually driven diagnostic domains in medicine.

P. Dhaimade, R. Henderson · 0 citations
Open access Aug 2026

Evaluating large language models in patient education: a comparative analysis addressing frequently asked questions in peri-acetabular osteotomy.

There may be significant differences in how effectively LLMs support patients with surgical queries, particularly in areas needing detailed explanation, and usually required minimal clarification in areas needing detailed explanation.

T. Davis, B. Guevel, K. Logishetty et al. · 0 citations
Open access Aug 2026

Comparative Evaluation of Large Language Models as Virtual Orthodontic Patient Information Assistants

It is demonstrated that LLMs can serve as auxiliary tools in providing orthodontic patient information but responses should be checked by an expert before being presented to patients and should be adapted into simpler and more understandable language.

Alperen Erdoğan, Orhan Çiçek · 0 citations
Open access Sep 2026

Comparative Evaluation of Large Language Model Interfaces in Third Molar Surgery Complication Scenarios: Response Quality, Clinical Content, Potential Clinical Risk, Readability, and Externally Observable Response Latency—A Cross-Sectional Comparative Benchmark Study

In this single-generation benchmark, the sampled outputs showed different profiles in specialist-rated quality, clinical content, potential clinical risk, readability, and externally observable response latency under the specific interface configurations tested.

I. Karaca, Esmanur Başer, Emre Ulubaş et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.