Evaluating the informational accuracy of large language models in patient‑directed orthodontic retainer guidance: a cross‑sectional comparison of Chat GPT‑4.1, Gemini 2.5, Microsoft Copilot GPT-4.1 and DeepSeek‑V3.
Findings suggest that AI models such as ChatGPT and Gemini can generate patient‑directed orthodontic retainer information with high informational accuracy under controlled evaluation conditions, however, specialist oversight remains essential to ensure clinical applicability.
Abstract
INTRODUCTION
The use of artificial intelligence (AI) in orthodontic practice is increasing rapidly; however, there is a notable lack of research evaluating the accuracy of large language models (LLMs) in educating patients about orthodontic retainers and related guidelines.
Materials And Methods
This study utilized a cross-sectional, repeated‑measures comparative evaluation design after receiving exemption from the institutional ethics committee. A set of 110 questions related to orthodontic retainers was compiled from previous articles addressing concerns about retainers and approved by a panel of three orthodontists. These questions were submitted to large language models (LLMs), including ChatGPT, Copilot, DeepSeek and Google Gemini. The responses were then reviewed by six independent orthodontists, who rated them using a modified five-point Likert scale.
Results
The overall accuracy revealed that 68.6% of responses scored 4, while 15.3% achieved a perfect score of 5. Among the LLMs, Gemini ranked first with 96.8%, closely followed by ChatGPT at 95.6%, indicating comparable high‑level performance between these models, while DeepSeek (76.9%) and Copilot (66.2%) demonstrated comparatively lower accuracy. Gemini produced a higher proportion of perfect scores, whereas ChatGPT consistently achieved strong ratings. The mean ratings across six raters demonstrated strong reliability (ICC = 0.81), reflecting expert agreement.
Conclusions
Findings suggest that AI models such as ChatGPT and Gemini can generate patient‑directed orthodontic retainer information with high informational accuracy under controlled evaluation conditions. However, specialist oversight remains essential to ensure clinical applicability. Future research using larger and more diverse datasets is needed to assess broader educational and communication‑related outcomes.
It is demonstrated that LLMs can serve as auxiliary tools in providing orthodontic patient information but responses should be checked by an expert before being presented to patients and should be adapted into simpler and more understandable language.
Alperen Erdoğan, Orhan Çiçek· Journal of International Den...· 0 citations
While LLMs show strong potential in supporting dental education through standardised exams, their performance varies by model and question type, and further improvements are needed to enhance reliability across different dental disciplines.
N. Acar, Fatih Sengul, Periş Çelikel et al.· European journal of dental e...· 0 citations
G Gemini demonstrated superior clinical precision and safety, whereas Claude provided more comprehensive and readable explanations, which support the integration of LLMs as pragmatic, high-ecological-validity complementary tools for patient education, while emphasizing the persistent necessity for professional clinical...
Ramazan Ağırağaç, Vedat Yüksekkaya· Journal of Stomatology Oral...· 0 citations
Objective: Although the integration of large language models (LLMs) into dental education is rapidly increasing, their actual performance in domain-specific assessments remains unclear. This study aimed to evaluate and compare the accuracy of four LLMs (ChatGPT-4.0, Gemini Advanced 1.5 Pro, DeepSeek-V3, and Perplexity)...
AI chatbots can generate information with potential clinical relevance in anterior implant dentistry; however, variability in informational reliability persists and expert supervision remains essential before integrating such tools into clinical education.
Dalndushe Abdulai, Raghıb Suradı, Mehran Moghbel· Journal of Health Sciences a...· 0 citations
This study compared the performance of three artificial intelligence–based chatbots (ChatGPT-4o, ChatGPT-4.5, and Gemini 2.5 Pro) on orthodontic questions from the Turkish Dental Specialty Examination at two testing time points. A total of 179 orthodontic multiple-choice questions from 18 examinations conducted between...