Skip to content
Review Open access

Evaluating large language models as tools for public health education on scoliosis

Jul 2026 · Frontiers in Public Health · Vol 14 · 0 citations · 27 references
Medicine

TL;DR

Although all three LLMs achieved favorable overall ratings, Claude performed significantly better than ChatGPT and DeepSeek for scoliosis FAQs taking into account its higher accuracy and consistency.

Abstract

Purpose This study was aimed to compare the efficacy of three most popular large language models (LLMs)—Claude Opus 4.6, ChatGPT Thinking 5.4 and DeepSeek v3.2 in answering frequently asked questions (FAQs) about scoliosis. Methods 20 scoliosis related questions (four categories, five questions in each category) were submitted to each LLM. A panel of 9 experts (two spine surgeons, two pediatric orthopedic surgeons and five physical therapists, all blinded to the LLMs and responses) rated independently each response generated by LLMs on a 6 points Likert scale (1 as strongly disagree to 6 as strongly agree). 540 total ratings were collected. Intergroup comparisons were conducted by Kruskal Wallis test and Mann Whitney U pairwise tests. Paired question level analysis was achieved by Friedman test and Wilcoxon signed rank comparisons. Results Claude’s score was 5.53 ± 0.76 much higher than both ChatGPT (4.84 ± 0.86, p < 0.001) and DeepSeek (4.86 ± 0.84, p < 0.001), but no difference was found between ChatGPT and DeepSeek (p = 0.749). Claude performed on top for 19 of 20 questions (95%) and was favored by 7 of 9 reviewers. Consistency of Claude was also highest [CV = 13.8% vs. 17.9% (ChatGPT) and 17.2% (DeepSeek)]. Conclusion Although all three LLMs achieved favorable overall ratings (>4.8/6), Claude performed significantly better than ChatGPT and DeepSeek for scoliosis FAQs taking into account its higher accuracy and consistency. Within the scope of the present evaluation, Claude demonstrated the strongest overall performance among the three LLMs tested.

Read PDF

Similar papers

Open access Aug 2026

Evaluating large language models in patient education: a comparative analysis addressing frequently asked questions in peri-acetabular osteotomy.

There may be significant differences in how effectively LLMs support patients with surgical queries, particularly in areas needing detailed explanation, and usually required minimal clarification in areas needing detailed explanation.

T. Davis, B. Guevel, K. Logishetty et al. · 0 citations
Open access Aug 2026

Accuracy and Consistency of Three Large Language Models on Fixed Prosthodontics Questions

Large language models (LLMs) are increasingly used in dental education and clinical settings, but their accuracy and consistency in fixed prosthodontics remain uncertain. Objectives: To compare the accuracy (using a strict three-attempt criterion) and repeated response consistency of ChatGPT-5.4, Gemini 3.1 Pro, and Cl...

A. Bashir, Moeen Ud Din Ahmad, Ussamah Waheed Jatala et al. · 0 citations
Review Open access Aug 2026

Vocal Cord Dysfunction: Evaluating the Utility of AI Large Language Models for Patient Education

ABSTRACT Objective To compare provider preferences for patient education materials generated by OpenEvidence, ChatGPT 5 Extended Thinking, and a laryngologist with fellowship training in response to common patient questions about vocal cord dysfunction (VCD). Study Design Cross‐sectional survey study. Methods A laryngo...

Kyle Cook, Phil Tseng, David Ahmadian et al. · 0 citations
Open access Sep 2026

Evidence on Demand: OpenEvidence versus Established LLMs on the Orthopaedic In-Training Examination

Large language models (LLMs) are increasingly used in medical education, yet head-to-head comparisons of contemporary multimodal and clinically oriented (retrieval-grounded) models on orthopedic in-service examinations remain limited. A cross-sectional comparative study was conducted using 265 multiple-choic...

Asim A. Khan, S. Lalvani, Sam Pourarbab et al. · 0 citations
Review Open access Aug 2026

Evaluation and Comparison of Large Language Model Responses to Frequently Asked Questions Regarding Patellofemoral Pain Syndrome: A Quality and Readability Assessment Study

The evaluated LLMs showed small and method-dependent differences in DISCERN-based information quality and variable differences in readability, but the findings should not be interpreted as evidence of factual accuracy, clinical safety, or suitability for individualized decision-making.

Oktay Polat, Berk Koncalıoğlu, Mert Gündoğdu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.