Skip to content
Open access

Performance comparison of large language models in interpreting clinical guidelines for migraine prevention: A multidimensional analysis

Sep 2026 · Neurology Asia · 0 citations

Abstract

Background & Objective: Large language models (LLMs) such as DeepSeek-R1, Gemini-2.5 Pro, ChatGPT-5 Thinking, and Grok-4 Expert are increasingly applied in medical contexts, yet their reliability in evidence-based clinical domains like migraine prophylaxis remains uncertain. This study aimed to compare the performance of four leading LLMs in interpreting and applying the International Headache Society’s global practice recommendations for preventive pharmacological treatment of migraine. Methods: Sixteen standardized clinical scenario questions derived from the IHS guideline were presented identically to each model. Responses were evaluated by blinded expert raters across five dimensions—Accuracy, Overconclusiveness, Supplementary Value, Incompleteness, and Readability—using a 10-point Likert scale. Readability was further analyzed using composite indices from readabilityformulas.com. Results: No significant inter-model differences were observed in Accuracy (P = 0.856), Overconclusiveness (P = 0.400), or Incompleteness (P = 0.531). However, DeepSeek-R1 provided significantly more Supplementary Information than Gemini-2.5 Pro (P = 0.010) and Grok-4 Expert (P = 0.030). Readability analysis further revealed substantial variation across models (P < 0.001), with DeepSeek-R1 generating the most accessible outputs. Conclusion: While all four models exhibited comparable adherence to guideline-based content, DeepSeek-R1 demonstrated superior performance in supplementary informational value and readability. These findings highlight the importance of evaluating LLMs not only for accuracy but also for their capacity to enhance clinical communication and decision support.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.