Skip to content
Review Open access

Artificial Intelligence Answering Patient Questions About Developmental Dysplasia of the Hip: Accuracy, Readability, and Clinical Utility.

Aug 2026 · Military Medicine · 0 citations
Medicine

TL;DR

It is suggested that AI chatbots may serve as useful adjunct educational tools for patients and caregivers within military healthcare systems, however, variability in expert interpretation and the potential for patient misinterpretation underscore the continued importance of clinician oversight when integrating AI-generated information into patient education.

Abstract

INTRODUCTION Artificial intelligence (AI) tools are increasingly used by patients and caregivers seeking medical information. Developmental dysplasia of the hip (DDH) is a common pediatric condition that frequently prompts parental questions regarding screening, diagnosis, and treatment. The purpose of this study was to evaluate the accuracy, completeness, and clarity of responses generated by a Department of Defense-approved AI chatbot (NIPR-GPT) when answering commonly asked DDH-related questions.

Methods

Twelve frequently asked patient questions related to DDH were submitted to the NIPR-GPT chatbot. Each question was entered into a new chatbot session to minimize contextual bias. AI-generated responses were independently evaluated by 5 fellowship-trained pediatric orthopedic surgeons. Responses were graded using a 4-point scale assessing accuracy, completeness, and clarity: (1) unsatisfactory (major inaccuracies requiring substantial correction), (2) satisfactory with moderate clarification required, (3) satisfactory with minimal clarification required, and (4) excellent with no clarification required. Inter-rater reliability was assessed using intraclass correlation coefficients (ICC) and weighted kappa statistics, while overall internal consistency among raters was assessed using Krippendorff's alpha.

Results

All AI-generated responses (100%) were rated either satisfactory or excellent. Six responses (50%) were graded excellent, requiring no clarification, and 6 responses (50%) were graded satisfactory with minimal clarification required. No responses were graded as unsatisfactory or requiring major correction. Single-rater reliability among individual reviewers was poor (ICC [2, 1] = 0.12), reflecting variability in reviewer thresholds. However, reliability improved when scores were averaged across raters, demonstrating moderate agreement (ICC [2, k] = 0.41). Internal consistency among reviewers was Krippendorff's α = 0.163, indicating heterogeneity in reviewer grading but not systematic deficiencies in AI responses. Five fellowship-trained pediatric orthopedic surgeons independently graded all 12 responses. No reviewer graded any response as unsatisfactory or requiring substantial correction. Mean question scores ranged from 2.8 to 4.0 across the 12 questions, with an overall mean score of approximately 3.6.

Conclusion

NIPR-GPT generated generally accurate, clear, and clinically appropriate responses to common DDH-related questions. Half of responses required no clarification, and none required major correction. These findings suggest that AI chatbots may serve as useful adjunct educational tools for patients and caregivers within military healthcare systems. However, variability in expert interpretation and the potential for patient misinterpretation underscore the continued importance of clinician oversight when integrating AI-generated information into patient education.

Read PDF

Similar papers

Review Open access Aug 2026

Accuracy and Readability of Generative Artificial Intelligence for Vascular Surgery Patients: A Specialist Based Evaluation Highlighting the Current Landscape of Safety Risks and Accessibility Gaps.

OBJECTIVE Patients frequently seek health information and medical advice from chatbots instead of consulting their physicians or referring to credible patient education resources provided by medical societies. A study to evaluate the quality, readability, and clinical appropriateness of ChatGPT generated answers to com...

Mario D'Oria, W. Dorigo, V. Alexiou et al. · 0 citations
Open access Sep 2026

Artificial Intelligence in Bariatric Patient Education: A Multi-rater Evaluation of Reliability, Readability, and Clinical Validity of ChatGPT 5.2

ChatGPT 5.2 is a valuable AI–assisted chatbot that facilitates patient education by providing responses regarding sleeve gastrectomy that are generally accurate and acceptable, but the categorization of 4–16% of the responses as "Incorrect," the overall difficult readability levels, and the significant variability obse...

Furkan Türkoğlu, Elif Nur Gencer, Emre Erdoğan · 0 citations
Open access Aug 2026

Accuracy and response repeatability of three large language models on undergraduate operative dentistry multiple-choice questions

All three large language models demonstrated high accuracy on operative dentistry MCQs, suggesting their potential utility as a supplementary educational tool, however, the residual error rates in the answers warrant the cautious integration of LLMs into dental education.

J. M. Antony, Nikhil Harikrishnan, N. Jayasheelan · 0 citations
Review Open access Sep 2026

Assessing the accuracy and usability of artificial intelligence-based language models in responding to common periodontal patient questions.

BACKGROUND Artificial intelligence-powered large language models (LLMs) are increasingly used by patients seeking quick information regarding dental and medical problems. Despite their growing popularity, concerns remain regarding the accuracy, clarity, and clinical usefulness of LLM-generated responses. This study aim...

Sajad Jahantigh, R. Amid, A. Moscowchi et al. · 0 citations
Review Open access Jul 2026

Evaluating the reliability, quality, and readability of AI-generated patient education on hallux valgus: a comparative study of large language models

Patients with hallux valgus increasingly seek health information through consumer-facing artificial intelligence (AI)–driven patient education tools, particularly large language model–based conversational agents. Although these tools offer rapid and accessible responses, concerns remain regarding the reliability,...

A. Koluman, Ebru Aloğlu Çiftçi, Mehmet Utku Çiftçi et al. · 0 citations
Aug 2026

Quality of AI-Generated Patient Education for Pre- and Post-Operative Tracheostomy Care.

AI chatbots can generate accurate and comprehensive responses to common tracheostomy care questions, demonstrating potential to support patient education, but they continue to lack guaranteed, verifiable sourcing.

Keer Zhang, Lauran K. Evans, Desiree Delavary et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.