Sep 2026· Journal of Health Sciences and Medicine· 0 citations· 22 references
TL;DR
AI chatbots can generate information with potential clinical relevance in anterior implant dentistry; however, variability in informational reliability persists and expert supervision remains essential before integrating such tools into clinical education.
Abstract
Aims: This study evaluated the quality, accuracy, and readability of responses generated by four Artificial Intelligence (AI) chatbots based on large language models (LLMs): ChatGPT-5.2, Gemini 3, DeepSeek-V3.2, and Grok-4.1, when responding to expert-generated questions in anterior implant dentistry. The aim was to evaluate their potential role as educational and adjunct informational tools in treatment planning and esthetic zone management, while also examining the readability of the generated responses.Methods: Thirty-six standardized questions covering diagnosis, esthetic risk assessment, implant positioning, surgical planning, peri-implant soft tissue management, esthetic complications, and preventive strategies were developed by three prosthodontists experienced in implant dentistry. Responses from ChatGPT, Gemini, DeepSeek, and Grok were independently evaluated by three experts using the modified DISCERN (mDISCERN) score, Global Quality Score (GQS), and a five-point Accuracy Score. Misinformation and Harm scores were also recorded. Repeated-measures comparisons were performed using the Friedman test with Wilcoxon signed-rank post hoc analysis and Bonferroni correction.Results: ChatGPT demonstrated the highest mean scores for informational reliability (mDISCERN: 3.33±0.72) and accuracy (3.56±0.84); however, differences in accuracy were not statistically significant. Gemini showed the highest numerical GQS value (GQS: 3.75±0.77); however, the overall effect size was small, and no GQS pairwise comparison remained statistically significant after adjustment. Grok showed intermediate performance, whereas DeepSeek showed lower numerical reliability scores, with significant pairwise differences identified only in comparison with ChatGPT and Gemini. Significant differences were observed for mDISCERN [χ² (3)=17.52, p0.05). Readability analysis showed a small but statistically significant difference in FRES among the chatbots, with Gemini demonstrating easier readability than Grok after Bonferroni adjustment, whereas FKGL did not differ significantly among the models; overall, the responses required a high-school to early undergraduate reading level.Conclusion: AI chatbots can generate information with potential clinical relevance in anterior implant dentistry; however, variability in informational reliability persists. Although ChatGPT showed slightly higher numerical scores for some outcomes, the observed differences were generally small, and no chatbot demonstrated clear superiority across all evaluated measures. Expert supervision therefore remains essential before integrating such tools into clinical education.
In this single-generation benchmark, the sampled outputs showed different profiles in specialist-rated quality, clinical content, potential clinical risk, readability, and externally observable response latency under the specific interface configurations tested.
I. Karaca, Esmanur Başer, Emre Ulubaş et al.· Journal of Clinical Medicine· 0 citations
G Gemini demonstrated superior clinical precision and safety, whereas Claude provided more comprehensive and readable explanations, which support the integration of LLMs as pragmatic, high-ecological-validity complementary tools for patient education, while emphasizing the persistent necessity for professional clinical...
Ramazan Ağırağaç, Vedat Yüksekkaya· Journal of Stomatology Oral...· 0 citations
Objective: Although the integration of large language models (LLMs) into dental education is rapidly increasing, their actual performance in domain-specific assessments remains unclear. This study aimed to evaluate and compare the accuracy of four LLMs (ChatGPT-4.0, Gemini Advanced 1.5 Pro, DeepSeek-V3, and Perplexity)...
The findings describe informational quality, not clinical safety or decision-making, and support specialist verification before clinical use, and support specialist verification before clinical use.
Aims: This study aimed to identify the clinically relevant patient questions about apical root-end resection based on international consensus guidelines, and to systematically evaluate the accuracy, quality, usefulness, and readability of generated responses by four Artificial Intelligence (AI)-powered conversational a...
T. Paksoy, Selin Gaş, Seval Ceylan Şen et al.· Journal of Dental Sciences a...· 0 citations
PURPOSE
To compare the information quality, accuracy, and readability of patient-directed responses generated by large language models (LLMs), including ChatGPT-o3, ChatGPT-5.2, Gemini 3, and DeepSeek, regarding robotic-assisted total knee arthroplasty (RA-TKA).
METHODS
Thirty frequently asked patient questions were...
U. Kolaç, Mazlum Veysel Sili, Orhan Mete Karademir et al.· Knee (Oxford)· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.