It is suggested that AI chatbots may serve as useful adjunct educational tools for patients and caregivers within military healthcare systems, however, variability in expert interpretation and the potential for patient misinterpretation underscore the continued importance of clinician oversight when integrating AI-generated information into patient education.
Abstract
INTRODUCTION
Artificial intelligence (AI) tools are increasingly used by patients and caregivers seeking medical information. Developmental dysplasia of the hip (DDH) is a common pediatric condition that frequently prompts parental questions regarding screening, diagnosis, and treatment. The purpose of this study was to evaluate the accuracy, completeness, and clarity of responses generated by a Department of Defense-approved AI chatbot (NIPR-GPT) when answering commonly asked DDH-related questions.
Methods
Twelve frequently asked patient questions related to DDH were submitted to the NIPR-GPT chatbot. Each question was entered into a new chatbot session to minimize contextual bias. AI-generated responses were independently evaluated by 5 fellowship-trained pediatric orthopedic surgeons. Responses were graded using a 4-point scale assessing accuracy, completeness, and clarity: (1) unsatisfactory (major inaccuracies requiring substantial correction), (2) satisfactory with moderate clarification required, (3) satisfactory with minimal clarification required, and (4) excellent with no clarification required. Inter-rater reliability was assessed using intraclass correlation coefficients (ICC) and weighted kappa statistics, while overall internal consistency among raters was assessed using Krippendorff's alpha.
Results
All AI-generated responses (100%) were rated either satisfactory or excellent. Six responses (50%) were graded excellent, requiring no clarification, and 6 responses (50%) were graded satisfactory with minimal clarification required. No responses were graded as unsatisfactory or requiring major correction. Single-rater reliability among individual reviewers was poor (ICC [2, 1] = 0.12), reflecting variability in reviewer thresholds. However, reliability improved when scores were averaged across raters, demonstrating moderate agreement (ICC [2, k] = 0.41). Internal consistency among reviewers was Krippendorff's α = 0.163, indicating heterogeneity in reviewer grading but not systematic deficiencies in AI responses. Five fellowship-trained pediatric orthopedic surgeons independently graded all 12 responses. No reviewer graded any response as unsatisfactory or requiring substantial correction. Mean question scores ranged from 2.8 to 4.0 across the 12 questions, with an overall mean score of approximately 3.6.
Conclusion
NIPR-GPT generated generally accurate, clear, and clinically appropriate responses to common DDH-related questions. Half of responses required no clarification, and none required major correction. These findings suggest that AI chatbots may serve as useful adjunct educational tools for patients and caregivers within military healthcare systems. However, variability in expert interpretation and the potential for patient misinterpretation underscore the continued importance of clinician oversight when integrating AI-generated information into patient education.
OBJECTIVE
Patients frequently seek health information and medical advice from chatbots instead of consulting their physicians or referring to credible patient education resources provided by medical societies. A study to evaluate the quality, readability, and clinical appropriateness of ChatGPT generated answers to com...
Mario D'Oria, W. Dorigo, V. Alexiou et al.· European Journal of Vascular...· 0 citations
ChatGPT 5.2 is a valuable AI–assisted chatbot that facilitates patient education by providing responses regarding sleeve gastrectomy that are generally accurate and acceptable, but the categorization of 4–16% of the responses as "Incorrect," the overall difficult readability levels, and the significant variability obse...
Furkan Türkoğlu, Elif Nur Gencer, Emre Erdoğan· Archives of Current Medical...· 0 citations
All three large language models demonstrated high accuracy on operative dentistry MCQs, suggesting their potential utility as a supplementary educational tool, however, the residual error rates in the answers warrant the cautious integration of LLMs into dental education.
J. M. Antony, Nikhil Harikrishnan, N. Jayasheelan· Frontiers in Dental Medicine· 0 citations
BACKGROUND
Artificial intelligence-powered large language models (LLMs) are increasingly used by patients seeking quick information regarding dental and medical problems. Despite their growing popularity, concerns remain regarding the accuracy, clarity, and clinical usefulness of LLM-generated responses. This study aim...
Sajad Jahantigh, R. Amid, A. Moscowchi et al.· Clinical Advances in Periodo...· 0 citations
Patients with hallux valgus increasingly seek health information through consumer-facing artificial intelligence (AI)–driven patient education tools, particularly large language model–based conversational agents. Although these tools offer rapid and accessible responses, concerns remain regarding the reliability,...
A. Koluman, Ebru Aloğlu Çiftçi, Mehmet Utku Çiftçi et al.· BMC Medical Informatics and...· 0 citations
AI chatbots can generate accurate and comprehensive responses to common tracheostomy care questions, demonstrating potential to support patient education, but they continue to lack guaranteed, verifiable sourcing.
Keer Zhang, Lauran K. Evans, Desiree Delavary et al.· Otolaryngology Head & Neck S...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.