Skip to content

Performance of large language models in electrocardiogram interpretation: A comparative study.

Jul 2026 · Journal of Electrocardiology · Vol 98, pp. 154411 · 0 citations · 34 references
Medicine

Abstract

Purpose

As large language models (LLMs) are increasingly used to interpret medical concerns, rigorous evaluation of their performance on clinically relevant tasks is essential. However, the new state-of-the-art models from OpenAI and Google as of February 2026, have not been evaluated for their accuracy and consistency in interpreting ECGs independent of clinical context. We aim to compare ChatGPT (GPT-5.2 Thinking) and Gemini (Gemini 3 Pro) on electrical axis and heart rhythm identification to assess current clinical usability and identify areas for improvement. METHODOLOGY ECGs were obtained from the Lobachevsky University Electrocardiography Databases on PhysioNet. The LLM responses were evaluated for first-shot accuracy and consistency across three different trials. First-shot accuracy was further split into the various types of rhythms and axes to examine systematic trends in model outputs.

Results

Both models demonstrated comparable overall first-shot accuracies for electric axis and rhythm classification. Both models struggled with less common rhythm categories, including multifocal rhythms and tachycardias. The Macro F1 analysis indicated low overall classification performance for both ChatGPT and Gemini in terms of axis and rhythm. ChatGPT achieved a higher Macro F1 point estimate for axis classification compared with Gemini, though both models struggled with a Macro F1 of 45.2% and 32.4% respectively.

Conclusion

The Macro F1 scores of both ChatGPT and Gemini suggest that they are not reliable for independent clinical diagnoses in cardiology, with difficulty shown in interpreting ECGs for rhythm and axis. This investigation aims to guide continued improvement of these LLM models for physician assistance.

View source