Benchmarking Large Language Models in Complex Hypersomnolence Disorders: A Comparative Clinical Analysis of ChatGPT and NotebookLM on Narcolepsy Guidelines
Jul 2026· The European Research Journal· pp. 1011-1022· 0 citations· 2 references
TL;DR
ChatGPT demonstrated superior accuracy and internal consistency in answering narcolepsy-related clinical questions compared with NotebookLM, however, neither model showed high proficiency in providing necessary additional clinical information.
Abstract
Objective: To compare the accuracy of 2 artificial intelligence models, ChatGPT and NotebookLM, in answering clinical questions regarding narcolepsy management.
Methods: A set of 30 clinical questions was developed based on 2 reference documents: the 2021 European guideline on the management of narcolepsy in adults and children, and the American Academy of Sleep Medicine clinical practice guideline for the treatment of central disorders of hypersomnolence. Both models were queried with the question set. Three independent scorers evaluated the responses across 4 categories (Accuracy, Evidence Reasoning, Additional Information, and Information Integration) using a binary scale (1 = criterion met, 0 = criterion not met). Interrater reliability was assessed using Cohen kappa and Fleiss kappa. Performance comparisons between the models were analyzed using the McNemar test, Wilcoxon signed-rank test, and Mann-Whitney U test, while scorer consistency was checked via the Friedman test.
Results: ChatGPT achieved 261 of 360 positive ratings (72.5%), whereas NotebookLM achieved 179 of 360 (49.7%). ChatGPT scored significantly higher than NotebookLM in Accuracy (P < .001) and Evidence Reasoning (P < .001). No statistically significant differences were observed between the 2 models in Additional Information or Information Integration.
Conclusion: ChatGPT demonstrated superior accuracy and internal consistency in answering narcolepsy-related clinical questions compared with NotebookLM. However, neither model showed high proficiency in providing necessary additional clinical information.
Background & Objective: Large language models (LLMs) such as DeepSeek-R1, Gemini-2.5 Pro, ChatGPT-5 Thinking, and Grok-4 Expert are increasingly applied in medical contexts, yet their reliability in evidence-based clinical domains like migraine prophylaxis remains uncertain. This study aimed to compare the performance...
Li Xu, Xu Qiu, Jia-Yi Deng et al.· Neurology Asia· 0 citations
The findings primarily highlight the need for expanded evaluation with multiple expert raters, patient participants, and more robust quality assessment methods before considering any practical or clinical use of LLMs.
Praghya Godavarthy, Manish Ravi Iyer, Ravinraj Munnavan et al.· Journal of Neurogastroentero...· 0 citations
Background: Artificial intelligence (AI) integration offers significant potential to improve mental healthcare, however, the reliability of large language models (LLMs) in performing nuanced clinical tasks remains an important and largely unanswered question. This study aimed to evaluate ChatGPT’s performance in scorin...
BackgroundLipedema is a frequently misdiagnosed chronic condition that significantly impacts patients' quality of life. As artificial intelligence (AI)-based large language models (LLMs) become increasingly integrated into healthcare communication, their accuracy and consistency in providing patient-centered informatio...
Rabia Sanır, E. Türkmen, E. Giray et al.· Phlebology· 0 citations
Although the four LLMs generally provide accurate and pertinent information regarding type 2 diabetes, enduring limits in actionability and inconsistencies among models in content completeness and understandability restrict their effective use in diabetic patient education.
Zhaoxia Huang, Yuxin Bai, Jun-Yue Luo et al.· Frontiers in Public Health· 0 citations
MIND was able to answer many questions about escitalopram in a manner deemed accurate, complete, and safe by psychiatrists the majority of the time, and represents a step towards building safe LLM systems to enhance patient education in psychiatry.
Alexander J. Hish, A. Nagendran, S. Compton· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.