Skip to content
Open access

Evaluation of large language model responses to expert questions in anterior implant dentistry: quality, accuracy, and readability

Sep 2026 · Journal of Health Sciences and Medicine · 0 citations · 22 references

TL;DR

AI chatbots can generate information with potential clinical relevance in anterior implant dentistry; however, variability in informational reliability persists and expert supervision remains essential before integrating such tools into clinical education.

Abstract

Aims: This study evaluated the quality, accuracy, and readability of responses generated by four Artificial Intelligence (AI) chatbots based on large language models (LLMs): ChatGPT-5.2, Gemini 3, DeepSeek-V3.2, and Grok-4.1, when responding to expert-generated questions in anterior implant dentistry. The aim was to evaluate their potential role as educational and adjunct informational tools in treatment planning and esthetic zone management, while also examining the readability of the generated responses.Methods: Thirty-six standardized questions covering diagnosis, esthetic risk assessment, implant positioning, surgical planning, peri-implant soft tissue management, esthetic complications, and preventive strategies were developed by three prosthodontists experienced in implant dentistry. Responses from ChatGPT, Gemini, DeepSeek, and Grok were independently evaluated by three experts using the modified DISCERN (mDISCERN) score, Global Quality Score (GQS), and a five-point Accuracy Score. Misinformation and Harm scores were also recorded. Repeated-measures comparisons were performed using the Friedman test with Wilcoxon signed-rank post hoc analysis and Bonferroni correction.Results: ChatGPT demonstrated the highest mean scores for informational reliability (mDISCERN: 3.33±0.72) and accuracy (3.56±0.84); however, differences in accuracy were not statistically significant. Gemini showed the highest numerical GQS value (GQS: 3.75±0.77); however, the overall effect size was small, and no GQS pairwise comparison remained statistically significant after adjustment. Grok showed intermediate performance, whereas DeepSeek showed lower numerical reliability scores, with significant pairwise differences identified only in comparison with ChatGPT and Gemini. Significant differences were observed for mDISCERN [χ² (3)=17.52, p0.05). Readability analysis showed a small but statistically significant difference in FRES among the chatbots, with Gemini demonstrating easier readability than Grok after Bonferroni adjustment, whereas FKGL did not differ significantly among the models; overall, the responses required a high-school to early undergraduate reading level.Conclusion: AI chatbots can generate information with potential clinical relevance in anterior implant dentistry; however, variability in informational reliability persists. Although ChatGPT showed slightly higher numerical scores for some outcomes, the observed differences were generally small, and no chatbot demonstrated clear superiority across all evaluated measures. Expert supervision therefore remains essential before integrating such tools into clinical education.

Read PDF

Similar papers

Open access Sep 2026

Comparative Evaluation of Large Language Model Interfaces in Third Molar Surgery Complication Scenarios: Response Quality, Clinical Content, Potential Clinical Risk, Readability, and Externally Observable Response Latency—A Cross-Sectional Comparative Benchmark Study

In this single-generation benchmark, the sampled outputs showed different profiles in specialist-rated quality, clinical content, potential clinical risk, readability, and externally observable response latency under the specific interface configurations tested.

I. Karaca, Esmanur Başer, Emre Ulubaş et al. · 0 citations
Aug 2026

Comparative Evaluation of Large Language Models in Answering Patient Questions Following Periodontal and Peri-Implant Examination: An Expert-Based Study.

G Gemini demonstrated superior clinical precision and safety, whereas Claude provided more comprehensive and readable explanations, which support the integration of LLMs as pragmatic, high-ecological-validity complementary tools for patient education, while emphasizing the persistent necessity for professional clinical...

Ramazan Ağırağaç, Vedat Yüksekkaya · 0 citations
Open access Sep 2026

Comparative Evaluation of Large Language Models’ Accuracy in Answering Multiple-Choice Restorative Dentistry Questions From a National Specialty Examination

Objective: Although the integration of large language models (LLMs) into dental education is rapidly increasing, their actual performance in domain-specific assessments remains unclear. This study aimed to evaluate and compare the accuracy of four LLMs (ChatGPT-4.0, Gemini Advanced 1.5 Pro, DeepSeek-V3, and Perplexity)...

Çilem Bulut, Gulben Colak, Gürkan Çolak · 0 citations
Sep 2026

Guideline-Based Evaluation of Five Generative Artificial Intelligence Chatbot Platforms for Clinician-Oriented Questions on Open Temporomandibular Joint Surgery.

The findings describe informational quality, not clinical safety or decision-making, and support specialist verification before clinical use, and support specialist verification before clinical use.

Selin Gaş, Erdinç Sulukan, Büşra Korkmaz · 0 citations
Open access Aug 2026

A multidisciplinary evaluation of AI-powered chatbots on apical root-end resection: assessing alignment with international endodontic guidelines-a comparative methodological study

Aims: This study aimed to identify the clinically relevant patient questions about apical root-end resection based on international consensus guidelines, and to systematically evaluate the accuracy, quality, usefulness, and readability of generated responses by four Artificial Intelligence (AI)-powered conversational a...

T. Paksoy, Selin Gaş, Seval Ceylan Şen et al. · 0 citations
Open access Sep 2026

Comparative quality, accuracy, and readability of large language model responses to patient questions about robotic-assisted total knee arthroplasty.

PURPOSE To compare the information quality, accuracy, and readability of patient-directed responses generated by large language models (LLMs), including ChatGPT-o3, ChatGPT-5.2, Gemini 3, and DeepSeek, regarding robotic-assisted total knee arthroplasty (RA-TKA). METHODS Thirty frequently asked patient questions were...

U. Kolaç, Mazlum Veysel Sili, Orhan Mete Karademir et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.