Skip to content
Open access

Comparison of responses from large language models using artificial intelligence for parent-focused inquiries on clubfoot treatment and Ponseti management

Jul 2026 · Journal of Children's Orthopaedics · Vol 20, pp. 576 - 582 · 0 citations · 28 references

TL;DR

Google Gemini outperformed ChatGPT-4.0 and CoPilot in accuracy and comprehensiveness for ICTEV management using the Ponseti method, whereas ChatGPT-4.0 offered more readable but less detailed answers.

Abstract

Purpose: This study aimed to compare the accuracy, completeness, and readability of responses generated by three large language models (LLMs)—ChatGPT-4.0 (OpenAI), Microsoft CoPilot, and Google Gemini—regarding the treatment and management of idiopathic congenital talipes equinovarus (ICTEV) using the Ponseti method. Methods: Fifteen frequently asked questions were selected from pediatric orthopedic center websites, Google Trends analysis, and clinical experience. Each question was submitted verbatim in a new session to the three LLMs within 24 h. Seven board-certified pediatric orthopedic surgeons, blinded to the source, rated responses for accuracy (5-point Likert scale) and completeness (3-point Likert scale). Readability was assessed using the Flesch–Kincaid grade level. Mean scores ± standard deviation were calculated, and inter-rater reliability was estimated using the intraclass correlation coefficient (ICC). Group differences were tested with ANOVA and chi-squared tests (p < 0.05). Results: A total of 45 responses were evaluated. Gemini achieved the highest mean accuracy (4.1 ± 0.8), followed by CoPilot (3.6 ± 0.8) and ChatGPT-4.0 (3.3 ± 0.9), with significant differences among models (p < 0.001). Completeness ratings also differed significantly (p < 0.001). Readability analysis showed that ChatGPT produced shorter, more readable text, while Gemini generated longer, more complex responses. Inter-rater reliability was substantial for accuracy (ICC 0.715) and completeness (ICC 0.710). Conclusions: Google Gemini outperformed ChatGPT-4.0 and CoPilot in accuracy and comprehensiveness for ICTEV management using the Ponseti method. However, its complex responses may limit accessibility, whereas ChatGPT-4.0 offered more readable but less detailed answers. Level of Evidence: IV

Read PDF

Similar papers

Review Open access Aug 2026

Evaluation and Comparison of Large Language Model Responses to Frequently Asked Questions Regarding Patellofemoral Pain Syndrome: A Quality and Readability Assessment Study

Background: Patellofemoral pain syndrome (PFPS) is a common cause of anterior knee pain, and patients increasingly use large language models (LLMs) to obtain general medical information. However, the quality, reliability, and readability of LLM-generated responses to patient-oriented questions regarding PFPS remain unc...

Oktay Polat, Berk Koncalıoğlu, Mert Gündoğdu et al. · 0 citations
Sep 2026

Guideline-Based Evaluation of Five Generative Artificial Intelligence Chatbot Platforms for Clinician-Oriented Questions on Open Temporomandibular Joint Surgery.

OBJECTIVE To compare the quality and readability of responses from five generative artificial intelligence chatbot platforms to clinician-oriented questions on open temporomandibular joint (TMJ) surgery against guideline-based reference answers. MATERIAL AND METHODS Forty questions across eight domains were submitted...

Selin Gaş, Erdinç Sulukan, Büşra Korkmaz · 0 citations
Open access Aug 2026

Accuracy and Completeness of Contemporary Large Language Models in Prosthodontics: An Expert-Based Comparative Study

Purpose: To compare the scientific accuracy and response completeness of four contemporary LLMs: ChatGPT-5.2, Gemini 3, Copilot, and DeepSeek-V3.2 across major prosthodontic domains and question formats.Methods: Fifty prosthodontic questions covering removable prosthodontics, fixed prosthodontics, implantology, temporo...

Elif Yiğit İren, Hatice Betül Üçkuyu · 0 citations
Open access Aug 2026

How Accurate Are Public Perceptions? A Comparative Analysis of Artificial Intelligence Responses on Physical Therapy for Gonarthrosis

Purpose: This study aimed to compare the content quality, reliability, readability, and structure of responses from different artificial intelligence (AI) models to the most frequently asked public questions on physical therapy for gonarthrosis. Methods: Ten highly frequent questions were collected from Google Trends...

Bilgehan Kolutek Ay, Mustafa Tuna · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.