Skip to content
Open access

Benchmarking Large Language Models in Complex Hypersomnolence Disorders: A Comparative Clinical Analysis of ChatGPT and NotebookLM on Narcolepsy Guidelines

Jul 2026 · The European Research Journal · pp. 1011-1022 · 0 citations · 2 references

TL;DR

ChatGPT demonstrated superior accuracy and internal consistency in answering narcolepsy-related clinical questions compared with NotebookLM, however, neither model showed high proficiency in providing necessary additional clinical information.

Abstract

Objective: To compare the accuracy of 2 artificial intelligence models, ChatGPT and NotebookLM, in answering clinical questions regarding narcolepsy management. Methods: A set of 30 clinical questions was developed based on 2 reference documents: the 2021 European guideline on the management of narcolepsy in adults and children, and the American Academy of Sleep Medicine clinical practice guideline for the treatment of central disorders of hypersomnolence. Both models were queried with the question set. Three independent scorers evaluated the responses across 4 categories (Accuracy, Evidence Reasoning, Additional Information, and Information Integration) using a binary scale (1 = criterion met, 0 = criterion not met). Interrater reliability was assessed using Cohen kappa and Fleiss kappa. Performance comparisons between the models were analyzed using the McNemar test, Wilcoxon signed-rank test, and Mann-Whitney U test, while scorer consistency was checked via the Friedman test. Results: ChatGPT achieved 261 of 360 positive ratings (72.5%), whereas NotebookLM achieved 179 of 360 (49.7%). ChatGPT scored significantly higher than NotebookLM in Accuracy (P < .001) and Evidence Reasoning (P < .001). No statistically significant differences were observed between the 2 models in Additional Information or Information Integration. Conclusion: ChatGPT demonstrated superior accuracy and internal consistency in answering narcolepsy-related clinical questions compared with NotebookLM. However, neither model showed high proficiency in providing necessary additional clinical information.

Read PDF

Similar papers

Open access Sep 2026

Performance comparison of large language models in interpreting clinical guidelines for migraine prevention: A multidimensional analysis

Background & Objective: Large language models (LLMs) such as DeepSeek-R1, Gemini-2.5 Pro, ChatGPT-5 Thinking, and Grok-4 Expert are increasingly applied in medical contexts, yet their reliability in evidence-based clinical domains like migraine prophylaxis remains uncertain. This study aimed to compare the performance...

Li Xu, Xu Qiu, Jia-Yi Deng et al. · 0 citations
Review Open access Aug 2026

Comparative Analysis of Large Language Model Outputs for Pre-hypnotherapy Support in Irritable Bowel Syndrome - An Experimental Feasibility Study.

The findings primarily highlight the need for expanded evaluation with multiple expert raters, patient participants, and more robust quality assessment methods before considering any practical or clinical use of LLMs.

Praghya Godavarthy, Manish Ravi Iyer, Ravinraj Munnavan et al. · 0 citations
Review Open access Aug 2026

A Cautious Integration With AI in the Clinic: A Standardized-Patient Pilot Study of ChatGPT’s Reliability in Hamilton Depression Rating Scale Scoring

Background: Artificial intelligence (AI) integration offers significant potential to improve mental healthcare, however, the reliability of large language models (LLMs) in performing nuanced clinical tasks remains an important and largely unanswered question. This study aimed to evaluate ChatGPT’s performance in scorin...

Chun-Hung Chang, Szu-Wei Cheng, Wei-Jen Chen et al. · 0 citations
Sep 2026

Evaluation of the accuracy and reproducibility of large language models (ChatGPT, DeepSeek, Gemini) in responding to patient-centered lipedema questions.

BackgroundLipedema is a frequently misdiagnosed chronic condition that significantly impacts patients' quality of life. As artificial intelligence (AI)-based large language models (LLMs) become increasingly integrated into healthcare communication, their accuracy and consistency in providing patient-centered informatio...

Rabia Sanır, E. Türkmen, E. Giray et al. · 0 citations
Open access Aug 2026

Evaluating large language models using the Type 2 Diabetes Health Education guideline: a comparative analysis of ChatGPT-4.1, Claude-4.0, DeepSeek-V3, and ERNIE Bot 4.5 Turbo

Although the four LLMs generally provide accurate and pertinent information regarding type 2 diabetes, enduring limits in actionability and inconsistencies among models in content completeness and understandability restrict their effective use in diabetic patient education.

Zhaoxia Huang, Yuxin Bai, Jun-Yue Luo et al. · 0 citations
#small language model Preprint Aug 2026

Performance of a domain-specific large language model in answering patient questions in psychiatry

MIND was able to answer many questions about escitalopram in a manner deemed accurate, complete, and safe by psychiatrists the majority of the time, and represents a step towards building safe LLM systems to enhance patient education in psychiatry.

Alexander J. Hish, A. Nagendran, S. Compton · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.