Diagnostic accuracy of electronic medical record retrieval methods and a large language model for identifying cardiovascular events: a multisite retrospective validation study in a medical system in the United States
Aug 2026· BMJ Open· Vol 16, pp. e116133· 0 citations· 25 references
Medicine
TL;DR
In this multisite retrospective validation study, the LLM-assisted workflow showed strong but context-dependent performance for identifying cardiovascular events from the EMR, and ICD-based retrieval remained competitive for some use cases.
Abstract
Objective To compare the diagnostic accuracy of four available automated electronic medical record (EMR) retrieval methods, including a large language model (LLM)-assisted workflow, against manual chart adjudication for identifying cardiovascular events. Design Retrospective diagnostic accuracy study. Setting Three sites within a single US tertiary health system. Participants Two adult cohorts with previously adjudicated cardiovascular outcomes were included. Cohort 1 included 2258 patients treated with immune checkpoint inhibitors, and Cohort 2 included 1426 patients who underwent transcatheter aortic valve replacement. Primary and secondary outcome measures The reference standard was clinician manual chart adjudication. Outcomes included ischaemic stroke or transient ischaemic attack, myocardial infarction (MI), heart failure (HF) exacerbation or hospitalisation and a composite major adverse cardiovascular events (MACE) outcome. Automated retrieval methods included International Classification of Diseases (ICD) codes, primary diagnosis, problem list and a zero-shot LLM workflow. Area under the (receiver operating characteristic) curve (AUC), sensitivity, specificity and net reclassification improvement were assessed. Results In Cohort 1, the LLM achieved the highest AUC for stroke (0.920; 95% CI 0.881 to 0.958), MI (0.938; 95% CI 0.905 to 0.971) and composite MACE (0.880; 95% CI 0.854 to 0.907), whereas ICD-based retrieval had a higher AUC for HF (0.882; 95% CI 0.845 to 0.918 vs 0.873; 95% CI 0.831 to 0.914). In Cohort 2, the LLM achieved the highest AUC for all evaluated outcomes: stroke (0.915; 95% CI 0.862 to 0.968), MI (0.928; 95% CI 0.839 to 1.000), HF (0.844; 95% CI 0.803 to 0.884) and composite MACE (0.862; 95% CI 0.829 to 0.895). In Cohort 1, differences in AUC between the LLM and ICD methods were not statistically significant across outcomes, whereas in Cohort 2 the LLM showed significantly higher AUC for stroke and composite MACE. Conclusion In this multisite retrospective validation study, the LLM-assisted workflow showed strong but context-dependent performance for identifying cardiovascular events from the EMR. Performance varied by outcome and cohort, and ICD-based retrieval remained competitive for some use cases. These findings support a complementary role for LLM-assisted extraction in retrospective cardiovascular outcomes research.
ObjectiveTo assess the accuracy of codes and algorithms used to identify selected connective tissue diseases (CTDs) in electronic health records (EHRs) and administrative databases.
MethodsWe searched MEDLINE, Embase, and CENTRAL databases for studies that validated case definitions in EHRs against a reference standar...
C. Saka-Herrán, T. M. P. Kumar, D. Keane et al.· medRxiv· 0 citations
BACKGROUND
Accurate diagnostic coding of sepsis is essential for surveillance, resource allocation, and health policy planning. Studies assessing the usability of claims-based data (ICD-10 codes) compared to clinical criteria for sepsis surveillance are needed.
OBJECTIVES
To assess the concordance between ICD-10 diag...
P. Nauclér, S. D. van der Werff, Andreas Winroth et al.· Infectious Diseases· 0 citations
Abstract Objective To understand how prevalence estimates of single and multiple organ fibrosis have changed from 2012 to 2022. Design Retrospective population-based cohort study. Setting Data from primary (Clinical Practice Research Datalink Aurum) and secondary care (Hospital Episode Statistics Admitted Patient Care)...
G. M. Massen, Gisli R. Jenkins, R. Allen et al.· BMJ Open· 0 citations
Introduction Heart failure may be preceded by non-specific symptoms or indicators recorded in primary care, although these records do not necessarily represent heart failure or missed diagnosis. We mapped quantitative evidence on prediagnostic signals, investigation, referral, diagnosis setting and outcomes. Methods We...
W. Jerjes, A. El-Osta, A. Jhass et al.· Journal of Primary Care & Co...· 0 citations
Background Structured administrative fields in oncology electronic health records (EHRs) are used in studies but their interpretation may depend on when they are completed. We assessed whether dated primary care physician (PCP) documentation was associated with survival after accounting for documentation timing. Materi...
P. Heudel, J. Blay· ESMO real world data and dig...· 0 citations
Clinico is developed and evaluated, a coding agent that organises evidence by time and clinical relationships to adjudicate diagnoses and update their status during admission and had the highest observed estimates of principal-diagnosis agreement and complete code-set micro-F1 among evaluated methods.
Y. Li, Y. Shi, Y. Sun et al.· medRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.