Skip to content
Review Open access

Implementing LLMs in clinical practice from unstructured and semistructured electronic record data to analysis ready data sets in central nervous system tumors

Jul 2026 · Discover medicine · Vol 3 · 0 citations · 33 references

TL;DR

LLMs can accurately extract clinical features from EHRs with careful document selection and prompt design and were the most accurately extracted biomarkers across these cohorts, MGMT and IDH were the most accurately extracted biomarkers.

Abstract

Primary CNS tumors impact 25,000 individuals annually in the U.S., posing significant health challenges. Research is hindered by fragmented data across electronic health record (EHR) systems and the inefficiency of manual extraction for diagnostic, molecular, and treatment data. Large Language Models (LLMs) present a solution by automating extraction, enhancing efficiency and consistency, and enabling AI-driven insights. We analyzed 4,974 EHR documents from 256 patients with confirmed glioblastoma (GBM). All patients were treated on NCI NIH IRB (IRB00011862)-approved protocols 00-C-0074, 02C0064, 04C0200, 06C0112, 16-C-0081, and 20-C-0027. Cohort 1 (GBM, n = 109) was used to develop to the pipeline, while Cohorts 2 (various CNS histologies, n = 147) and 3 (GBM lesions, n = 15) served as testing sets. Clinical features extracted with GPT-4o via the NIH NIDAP Text Extraction Program (NTEP) included date of diagnosis, KPS, extent of resection, MGMT and IDH statuses, and radiation therapy start/end dates. Prompts underwent iterations, utilized JavaScript Object Notation (JSON) formatting, and outputs were then compared to the manual ground truth, established through detailed chart review of all 256 patients across three cohorts to extract key clinical variables of interest, allowing for the analysis of clinical feature and document type accuracy. Prompt refinement led to a 30-fold increase in prompt character count, achieving ≥ 95% accuracy for five of seven features in Cohort 1. MGMT accuracy increased from 26% to 99%, and radiation dates increased from 93% to 98%. KPS (82%) and extent of resection (84%) were less accurate. For Cohort 2, MGMT and IDH were extracted with the highest accuracy (87% and 90%, respectively), while Cohort 3 achieved slightly lower but still strong performance (70% and 78%). Across these cohorts, MGMT and IDH were the most accurately extracted biomarkers. Radiation therapy summaries were the most effective documents across all three cohorts based on the extraction rates of clinical features. LLMs can accurately extract clinical features from EHRs with careful document selection and prompt design. This method may support CNS tumor research and have broader clinical applications. Future work will expand feature sets and validate on external datasets.

Read PDF

Similar papers

Open access Sep 2026

Bridging the data latency gap: automated extraction of genomic biomarkers from unstructured clinical documents to support real-world oncology data

Real-world oncology data are essential for clinical research and precision cancer care. However, genomic biomarkers are often embedded in scanned, unstructured clinical documents requiring manual abstraction before becoming available in cancer registries, delaying real-world evidence generation. This study evaluated an...

Qian-Yun Luo, Rui Zhang, Nikitha Vobugari et al. · 0 citations
Open access Aug 2026

Incremental Diagnostic Value of Clinical Information for Large Language Models Across Multiple Organs: Retrospective Study

The addition of clinical information was associated with a numeric trend toward higher diagnostic accuracy overall, but this trend was heterogeneous across models and disease types, and no statistically significant improvement was demonstrated after adjustment for multiple comparisons.

Jin-Qi Zhang, Xiao-Yi Wang, Yanfeng Zhao et al. · 0 citations
Open access Aug 2026

Developing an open-source framework for LLM evaluation of patients using EHR clinical documentation; performance of LLMs relative to medical professionals

Current LLMs do not achieve inter-rater reliability levels comparable to medical professionals in clinical information extraction from ENT documentation, suggesting they are best suited for initial extraction with human verification rather than autonomous operation.

L. Barrett, N. Joshi, A. S. North et al. · 0 citations
Open access Sep 2026

Development and validation of a pragmatic pipeline for clinical free-text annotation using locally deployed open-weight large language models

Clinical information required for surgical data science (SDS) is frequently embedded in unstructured text. We developed and evaluated a reproducible pipeline for selecting locally deployed open-weight large language models (LLMs) for binary symptom annotation. In this retrospective single-center study, 1,100 German eme...

Jonas Henn, Alisa Stoll, P. Feodorovici et al. · 0 citations
Open access Sep 2026

Clinical Code Mapping with LLM Tool Use: A Pilot for Automated Data Extraction of Medication and Diagnosis Information from Unstructured Clinical Notes.

LLMs are suitable for information extraction of medications from clinical notes for use in research databases, however, for a clinical setting where the treatment of patients would be dependent on LLM performance, the current state-of-the-art open weight models are not accurate enough.

T. Spreuer, A. Günther, R. Majeed · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.