Skip to content
Open access

Benchmarking Large Language Models in Complex Hypersomnolence Disorders: A Comparative Clinical Analysis of ChatGPT and NotebookLM on Narcolepsy Guidelines

Jul 2026 · The European Research Journal · 0 citations · 2 references

Abstract

Objective: To compare the accuracy of 2 artificial intelligence models, ChatGPT and NotebookLM, in answering clinical questions regarding narcolepsy management. Methods: A set of 30 clinical questions was developed based on 2 reference documents: the 2021 European guideline on the management of narcolepsy in adults and children, and the American Academy of Sleep Medicine clinical practice guideline for the treatment of central disorders of hypersomnolence. Both models were queried with the question set. Three independent scorers evaluated the responses across 4 categories (Accuracy, Evidence Reasoning, Additional Information, and Information Integration) using a binary scale (1 = criterion met, 0 = criterion not met). Interrater reliability was assessed using Cohen kappa and Fleiss kappa. Performance comparisons between the models were analyzed using the McNemar test, Wilcoxon signed-rank test, and Mann-Whitney U test, while scorer consistency was checked via the Friedman test. Results: ChatGPT achieved 261 of 360 positive ratings (72.5%), whereas NotebookLM achieved 179 of 360 (49.7%). ChatGPT scored significantly higher than NotebookLM in Accuracy (P < .001) and Evidence Reasoning (P < .001). No statistically significant differences were observed between the 2 models in Additional Information or Information Integration. Conclusion: ChatGPT demonstrated superior accuracy and internal consistency in answering narcolepsy-related clinical questions compared with NotebookLM. However, neither model showed high proficiency in providing necessary additional clinical information.

Read PDF