Skip to content
Open access

A Multidimensional Evaluation of Large Language Model Responses to the 2024 ESC Hypertension Guidelines: A Comparative Study

Sep 2026 · Journal of Cardiovascular Development and Disease · 0 citations · 17 references

Abstract

Large language models (LLMs), including ChatGPT, Gemini, and DeepSeek, are increasingly used in medicine; however, their performance across multiple clinically relevant domains remains incompletely understood. This cross-sectional comparative study evaluated 225 responses generated by three LLMs to 75 clinical questions selected from the 2024 European Society of Cardiology (ESC) Hypertension Guidelines. Responses were independently assessed by two board-certified cardiologists across five predefined domains: accuracy, clinical relevance, completeness, absence of bias and misinformation, and consistency. No statistically significant differences were observed among the three models in accuracy, clinical relevance, completeness, or absence of bias and misinformation (all p > 0.05). A significant difference was identified in response consistency (p = 0.032), with post hoc analysis demonstrating a significant difference between ChatGPT and Gemini. Overall appropriateness scores did not differ significantly among the three LLMs (p = 0.227). These findings suggest that, although overall performance was comparable, response consistency represents an additional dimension that should be considered when evaluating LLMs for guideline-based clinical applications. Future studies incorporating broader clinical scenarios and updated LLM versions are warranted to further define their role in clinical decision support.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.