Skip to content
Open access

Response Quality of AI-Generated Answers to Orthodontic Patient Concerns Across Empathy, Accuracy, Comprehensiveness, Personalisation, Safety, and Patient-Centeredness: A Cross-Model, Bilingual Evaluation of Claude, GPT-4o, and Gemini

Oct 2026 · Healthcare · 0 citations · 24 references

Abstract

Background: Patients now consult large language models (LLMs) for orthodontic concerns, but most evaluations sit in English and focus on factual accuracy. We compared three current LLMs (Claude Opus 4.5, GPT-4o, Gemini 2.5 Flash) on patient-facing responses in English and Turkish, and tested whether model differences held up after adjusting for response length. Methods: Sixty standardised patient scenarios were prompted in both languages to each model via stateless single-turn API calls with identical token limits, yielding 360 responses. Three blinded orthodontists scored every response on empathy, accuracy, comprehensiveness, personalisation, safety, and patient-centeredness (1–5 Likert). Primary between-model contrasts used mixed-effects linear models with scenario as a random effect to accommodate the paired scenario structure, with length-adjusted specifications adding response word count as a covariate; one-way ANOVA and ANCOVA on the balanced English sample were reported concordantly as descriptive checks. Reliability was assessed with ICC(2,3) and Krippendorff’s α. Results: Forty-eight Turkish Gemini outputs (80%) were truncated mid-generation and were excluded, leaving 312 successfully completed responses. Between-model ranking is therefore reported on the balanced English cell (n = 60 per model). Unadjusted, Gemini scored highest on the English cell (total 4.04 vs. 3.86 GPT-4o, 3.85 Claude). But response length correlated strongly with scores (r = 0.75 total; r = 0.84 comprehensiveness), and mean Gemini English length (547 words) was roughly double Claude’s (265) and GPT-4o’s (288). After length adjustment, Gemini’s advantage on empathy, comprehensiveness, and personalisation disappeared (all p > 0.10), the total-score coefficient reversed (β = −0.065, p = 0.017), and Gemini scored lower on patient-centeredness (β = −0.153, p = 0.021). English scored substantially higher than Turkish on every dimension (3.92 vs. 3.51; d = 2.71), and patient-centeredness was the weakest dimension overall (3.27). Conclusions: Rubric-based scoring of LLM outputs is strongly associated with response length; unadjusted between-model comparisons may mainly reflect length differences rather than architectural differences. The English–Turkish gap and the Turkish Gemini truncation are robust to length adjustment. Per-language quality assurance—for both completeness and length-normalised content quality—is a prerequisite for patient-facing LLM deployment; English-only evaluation is not a safe proxy.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.