Skip to content
Review Open access

Comparative Analysis of Large Language Model Outputs for Pre-hypnotherapy Support in Irritable Bowel Syndrome - An Experimental Feasibility Study.

Aug 2026 · Journal of Neurogastroenterology and Motility · 0 citations
Medicine

TL;DR

The findings primarily highlight the need for expanded evaluation with multiple expert raters, patient participants, and more robust quality assessment methods before considering any practical or clinical use of LLMs.

Abstract

Background/Aims : Gut directed hypnotherapy can assist with irritable bowel syndrome, but access remains limited. This feasibility study conducted an exploratory comparison of 4 large language models (LLMs) to evaluate the readability and information quality of model generated responses to common patient questions about pre- hypnotherapy support. Methods : Four LLMs (ChatGPT-4o, Claude 3.7 Sonnet, Gemini, and DeepSeek-R1) were queried using a pre-specified set of 14 standardized patient-oriented prompts. Responses were assessed for readability using validated indices (FKGL, FRES, and CLI) and for information quality using the DISCERN instrument and 5 item Likert scale, rated by a single evaluator. Paired statistical tests were used, and results are reported with effect sizes and confidence intervals. Results : In this exploratory evaluation, DeepSeek-R1 and ChatGPT-4o generated responses with relatively higher scores across readability metrices (Flesch-Kincaid Grade Level, Flesch Reading Ease Score, and Coleman-Liau Index) and information quality measures (DISCERN and Likert), which are distinct constructs assessed by separate instruments. Claude 3.7 Sonnet produced denser with lower readability and reliability scores, and Gemini showed intermediate performance. Interpretation is limited using automated readability metrics and single-rate design. As the DISCERN and Likert scores were generated by a single rater, statistical comparisons involving these measures cannot be considered definitive. The results below are presented solely to illustrate observable trends within this exploratory dataset. Conclusion LLMs show potential for generating patient oriented pre-hypnotherapy information, but this feasibility study is not sufficient to determine clinical applicability. The findings primarily highlight the need for expanded evaluation with multiple expert raters, patient participants, and more robust quality assessment methods before considering any practical or clinical use.

Read PDF

Similar papers

Open access Aug 2026

Evaluating large language models using the Type 2 Diabetes Health Education guideline: a comparative analysis of ChatGPT-4.1, Claude-4.0, DeepSeek-V3, and ERNIE Bot 4.5 Turbo

Although the four LLMs generally provide accurate and pertinent information regarding type 2 diabetes, enduring limits in actionability and inconsistencies among models in content completeness and understandability restrict their effective use in diabetic patient education.

Zhaoxia Huang, Yuxin Bai, Jun-Yue Luo et al. · 0 citations
Sep 2026

Evaluation of the accuracy and reproducibility of large language models (ChatGPT, DeepSeek, Gemini) in responding to patient-centered lipedema questions.

BackgroundLipedema is a frequently misdiagnosed chronic condition that significantly impacts patients' quality of life. As artificial intelligence (AI)-based large language models (LLMs) become increasingly integrated into healthcare communication, their accuracy and consistency in providing patient-centered informatio...

Rabia Sanır, E. Türkmen, E. Giray et al. · 0 citations
Open access Sep 2026

Performance comparison of large language models in interpreting clinical guidelines for migraine prevention: A multidimensional analysis

Background & Objective: Large language models (LLMs) such as DeepSeek-R1, Gemini-2.5 Pro, ChatGPT-5 Thinking, and Grok-4 Expert are increasingly applied in medical contexts, yet their reliability in evidence-based clinical domains like migraine prophylaxis remains uncertain. This study aimed to compare the performance...

Li Xu, Xu Qiu, Jia-Yi Deng et al. · 0 citations
Open access Sep 2026

Performance of large language models in answering public questions about nutrition in cirrhosis: a comparative study.

BACKGROUND large language models (LLMs) are increasingly used for public health information, but their performance in nutrition advice for cirrhosis remains uncertain. We investigated four LLMs in answering public questions about cirrhosis nutrition across safety, accuracy, empathy, information reliability and quality,...

Jun-Zheng Li, Ying-Jie Wu, Man Yang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.