Skip to content
Open access

Comparative evaluation of large language models for patient-facing information in clear aligner therapy

Sep 2026 · APOS Trends in Orthodontics · 0 citations · 20 references

TL;DR

Large language models can provide appropriate patient-facing information about clear aligner therapy, though meaningful differences in communication quality and empathy exist, though meaningful differences in communication quality and empathy exist.

Abstract

Patients considering clear aligner therapy should understand treatment requirements, limitations, and the importance of professional supervision. As many patients now seek orthodontic information from artificial intelligence (AI) chatbots, this study evaluated how effectively these systems communicate patient-facing information related to clear aligner therapy. Twenty common patient-facing questions about clear aligner therapy were developed using conversational language. Three large language models (LLM), ChatGPT 5.2, Claude Sonnet 4.5, and Gemini Flash 3, generated responses using default settings. Three orthodontic specialists, blinded to model identity, rated responses for quality and empathy using 5-point Likert scales. Readability was assessed using the Flesch–Kincaid Grade Level. Inter-rater reliability was evaluated using intraclass correlation coefficients (2,1). Model performance was compared using repeated-measures analysis of variance with Bonferroni-adjusted post hoc testing. Inter-rater reliability ranged from fair to good. Significant differences were observed among models for quality (F[2,38] = 11.43; p < 0.001) and empathy (F[2,38] = 3.92; p = 0.028), while readability did not differ significantly ( p = 0.321). Gemini Flash 3 achieved the highest mean scores for both quality (4.57 ± 0.42) and empathy (4.28 ± 0.51). All models consistently encouraged professional orthodontic consultation and avoided recommending unsupervised treatment. Large language models can provide appropriate patient-facing information about clear aligner therapy, though meaningful differences in communication quality and empathy exist. Clinical oversight remains essential when integrating AI tools into orthodontic patient education.

Read PDF

Similar papers

Open access Aug 2026

Comparative Evaluation of Large Language Models as Virtual Orthodontic Patient Information Assistants

It is demonstrated that LLMs can serve as auxiliary tools in providing orthodontic patient information but responses should be checked by an expert before being presented to patients and should be adapted into simpler and more understandable language.

Alperen Erdoğan, Orhan Çiçek · 0 citations
Open access Oct 2026

Response Quality of AI-Generated Answers to Orthodontic Patient Concerns Across Empathy, Accuracy, Comprehensiveness, Personalisation, Safety, and Patient-Centeredness: A Cross-Model, Bilingual Evaluation of Claude, GPT-4o, and Gemini

Background: Patients now consult large language models (LLMs) for orthodontic concerns, but most evaluations sit in English and focus on factual accuracy. We compared three current LLMs (Claude Opus 4.5, GPT-4o, Gemini 2.5 Flash) on patient-facing responses in English and Turkish, and tested whether model differences h...

Mustafa Özcan, Ferdi Allaf · 0 citations
#small language model Review Open access Sep 2026

Comparative Evaluation of Large Language Models in Responding to Common Misconceptions Regarding Periodontal Disease and Oral Hygiene

Background/Objective: As large language models (LLMs) are increasingly used to obtain oral health information, their ability to provide accurate and comprehensive responses to patient-focused periodontal questions derived from commonly reported misconceptions is increasingly relevant to patient education. This study ai...

İsmail Gül, Resül Çolak, Merve Küçükoğlu Çolak et al. · 0 citations
Open access Aug 2026

Evaluating large language models in patient education: a comparative analysis addressing frequently asked questions in peri-acetabular osteotomy.

There may be significant differences in how effectively LLMs support patients with surgical queries, particularly in areas needing detailed explanation, and usually required minimal clarification in areas needing detailed explanation.

T. Davis, B. Guevel, K. Logishetty et al. · 0 citations
Aug 2026

Comparative Evaluation of Large Language Models in Answering Patient Questions Following Periodontal and Peri-Implant Examination: An Expert-Based Study.

G Gemini demonstrated superior clinical precision and safety, whereas Claude provided more comprehensive and readable explanations, which support the integration of LLMs as pragmatic, high-ecological-validity complementary tools for patient education, while emphasizing the persistent necessity for professional clinical...

Ramazan Ağırağaç, Vedat Yüksekkaya · 0 citations
Review Open access Aug 2026

Evaluation and Comparison of Large Language Model Responses to Frequently Asked Questions Regarding Patellofemoral Pain Syndrome: A Quality and Readability Assessment Study

The evaluated LLMs showed small and method-dependent differences in DISCERN-based information quality and variable differences in readability, but the findings should not be interpreted as evidence of factual accuracy, clinical safety, or suitability for individualized decision-making.

Oktay Polat, Berk Koncalıoğlu, Mert Gündoğdu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.