Comparison of responses from large language models using artificial intelligence for parent-focused inquiries on clubfoot treatment and Ponseti management
Jul 2026· Journal of Children's Orthopaedics· Vol 20, pp. 576 - 582· 0 citations· 28 references
TL;DR
Google Gemini outperformed ChatGPT-4.0 and CoPilot in accuracy and comprehensiveness for ICTEV management using the Ponseti method, whereas ChatGPT-4.0 offered more readable but less detailed answers.
Abstract
Purpose: This study aimed to compare the accuracy, completeness, and readability of responses generated by three large language models (LLMs)—ChatGPT-4.0 (OpenAI), Microsoft CoPilot, and Google Gemini—regarding the treatment and management of idiopathic congenital talipes equinovarus (ICTEV) using the Ponseti method. Methods: Fifteen frequently asked questions were selected from pediatric orthopedic center websites, Google Trends analysis, and clinical experience. Each question was submitted verbatim in a new session to the three LLMs within 24 h. Seven board-certified pediatric orthopedic surgeons, blinded to the source, rated responses for accuracy (5-point Likert scale) and completeness (3-point Likert scale). Readability was assessed using the Flesch–Kincaid grade level. Mean scores ± standard deviation were calculated, and inter-rater reliability was estimated using the intraclass correlation coefficient (ICC). Group differences were tested with ANOVA and chi-squared tests (p < 0.05). Results: A total of 45 responses were evaluated. Gemini achieved the highest mean accuracy (4.1 ± 0.8), followed by CoPilot (3.6 ± 0.8) and ChatGPT-4.0 (3.3 ± 0.9), with significant differences among models (p < 0.001). Completeness ratings also differed significantly (p < 0.001). Readability analysis showed that ChatGPT produced shorter, more readable text, while Gemini generated longer, more complex responses. Inter-rater reliability was substantial for accuracy (ICC 0.715) and completeness (ICC 0.710). Conclusions: Google Gemini outperformed ChatGPT-4.0 and CoPilot in accuracy and comprehensiveness for ICTEV management using the Ponseti method. However, its complex responses may limit accessibility, whereas ChatGPT-4.0 offered more readable but less detailed answers. Level of Evidence: IV
Background: Patellofemoral pain syndrome (PFPS) is a common cause of anterior knee pain, and patients increasingly use large language models (LLMs) to obtain general medical information. However, the quality, reliability, and readability of LLM-generated responses to patient-oriented questions regarding PFPS remain unc...
Oktay Polat, Berk Koncalıoğlu, Mert Gündoğdu et al.· Healthcare· 0 citations
OBJECTIVE
To compare the quality and readability of responses from five generative artificial intelligence chatbot platforms to clinician-oriented questions on open temporomandibular joint (TMJ) surgery against guideline-based reference answers.
MATERIAL AND METHODS
Forty questions across eight domains were submitted...
ChatGPT-4o and ChatGPT-5 provide generally satisfactory yet non-comprehensive, limited-quality information at a level above tenth-grade regarding hallux rigidus fusion surgery.
Kamil Balaban, Mehmet Batu Ertan, Mahmut Kalem· Digital Health· 0 citations
It is suggested that model selection influences expert-rated educational suitability and overall information quality, whereas topic complexity primarily affects readability.
Hao Wei, Sisi Sun, Ming-Xin Liu et al.· Journal of Visualized Experi...· 0 citations
Purpose: To compare the scientific accuracy and response completeness of four contemporary LLMs: ChatGPT-5.2, Gemini 3, Copilot, and DeepSeek-V3.2 across major prosthodontic domains and question formats.Methods: Fifty prosthodontic questions covering removable prosthodontics, fixed prosthodontics, implantology, temporo...
Purpose:
This study aimed to compare the content quality, reliability, readability, and structure of responses from different artificial intelligence (AI) models to the most frequently asked public questions on physical therapy for gonarthrosis. Methods: Ten highly frequent questions were collected from Google Trends...
Bilgehan Kolutek Ay, Mustafa Tuna· International Journal of Dig...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.