Artificial Intelligence in Bariatric Patient Education: A Multi-rater Evaluation of Reliability, Readability, and Clinical Validity of ChatGPT 5.2
Abstract
Background: Although artificial intelligence (AI) has been used in patient education for some time, the accuracy, reliability, and clinical appropriateness of AI-generated medical content remain inadequately defined and continue to be debated. This study aimed to evaluate the reliability, readability, and comprehensibility of responses generated by ChatGPT 5.2 to the most frequently asked questions related to sleeve gastrectomy, and to assess the responses by surgeons with varying levels of clinical expertise. Methods: Twenty-four questions regarding sleeve gastrectomy were asked to ChatGPT 5.2, and responses were evaluated using a Likert-Scale by three general surgeons with varying levels of expertise. The readability and comprehensibility analysis was also conducted, using Flesch–Kincaid Grade Level and Flesch Reading Ease Scores. Results: Although excellent intra-rater reliability was observed, inter-rater reliability was poor. The evaluator with the greatest clinical expertise assigned the highest mean score (2.92±0.776), whereas the evaluator possessing predominantly theoretical knowledge assigned the lowest mean score (2.58±0.881). In the readability analysis, responses in all subcategories were classified as “difficult to read,” whereas only the responses within the postoperative course category were categorized as “fairly difficult”. Conclusions: ChatGPT 5.2 is a valuable AI–assisted chatbot that facilitates patient education by providing responses regarding sleeve gastrectomy that are generally accurate and acceptable. Nevertheless, the categorization of 4–16% of the responses as "Incorrect," the overall difficult readability levels, and the significant variability observed among evaluators with different levels of clinical experience underscore the indispensable role of clinical expertise and specialized professional guidance in surgical practice.