Skip to content
#software testing Review Open access

Cutting chart-review time and improving database accuracy in inflammatory bowel disease with human-in-the-loop large language models

Aug 2026 · BMC Medical Informatics and Decision Making · 0 citations

TL;DR

An open-source, human-verified workflow using large language models can accelerate electronic health record abstraction while improving accuracy and supports broader adoption of transparent artificial intelligence methods in clinical research.

Abstract

Manual abstraction of electronic health records into research databases is a major bottleneck in clinical research, limiting scalability and introducing error. This challenge is particularly acute in inflammatory bowel disease surveillance, where clinically relevant variables are distributed across extensive longitudinal documentation. We evaluated whether a reproducible, open-source, human-in-the-loop workflow based on large language models could outperform standard manual chart review in both efficiency and accuracy. We developed a locally deployed, source-linked extraction pipeline using open-source large language models to recreate an existing inflammatory bowel disease surveillance database. The system employed a two-stage, activation-based architecture that extracted structured variables from clinical notes and returned each value with supporting source citations. In a controlled user study, four clinicians each annotated 20 patients: 10 by manual chart review and 10 by reviewing and correcting model-generated outputs using a custom, source-aware user interface. Primary outcomes were extraction time per patient and per variable, and extraction accuracy relative to ground truth. Time outcomes were compared using two-sided Mann–Whitney U tests due to non-normal distributions, with effect sizes and confidence intervals reported. Accuracy differences were summarized using absolute risk differences with confidence intervals. Median extraction time per patient decreased from 9.4 min with manual review to 3.6 min with model assistance, yielding a typical time saving of 5.27 min per patient and a large effect size ( r  = 0.747, 95% confidence interval 0.576–0.885). Extraction accuracy improved from 68% with manual abstraction to 89% with model-assisted annotation (risk difference 0.213; 95% confidence interval 0.102–0.316). Accuracy gains were greatest for variables requiring synthesis across multiple clinical notes, while performance on high-salience variables was comparable across workflows. An open-source, human-verified workflow using large language models can accelerate electronic health record abstraction while improving accuracy. By releasing both the extraction pipeline and user interface software, this study provides a reproducible and deployable template for scalable clinical database curation and supports broader adoption of transparent artificial intelligence methods in clinical research.

Read PDF

Similar papers

Review Jul 2026

Validation of a Text-Mining Tool for Extracting Routine Clinical Care Data in Early-Stage Resectable Non-Small Cell Lung Cancer.

PURPOSE Manual chart review (MR) of electronic health records (EHRs) is time-consuming, error-prone, and limits the reproducibility and scalability of real-world data (RWD) research. Automation and standardization using natural language processing (NLP) could improve efficiency and scalability. CTcue is an NLP-based software platform designed to extract structured and unstructured data from EHRs. This study evaluated the accuracy and efficiency of CTcue versus MR in patients with early-stage resectable non-small cell lung cancer (NSCLC). METHODS Included were all patients with stage I to III NSCLC who underwent lung resections between January 2018 and December 2021 at the Leiden University Medical Center, the Netherlands. Demographics, tumor characteristics, treatment, and outcomes were collected. CTcue performance was compared with MR using weighted F1-scores, accuracy, precision, and recall for categorical variables and Bland-Altman analysis for continuous variables. RESULTS Eighty-five patients (70.2% of patients from the manual cohort) were identified by both methods and included in the comparison. CTcue achieved weighted F1-scores >0.85 for seven of 15 categorical variables, including sex, tumor location, and deceased status, although some scores were based on low number of observations in both cohorts. Lower performance was observed for variables with varying terminology in documentation, such as Eastern Cooperative Oncology Group status and pathological N-stage. Continuous variables showed negligible mean differences, indicating good agreement. Survival outcomes were identical in both data sets. CONCLUSION CTcue performance for patient selection was lower than anticipated. However, it enables accurate, efficient extraction of structured and unstructured EHR data in early-stage NSCLC. Manual validation remains necessary for variables with varying terminology. Further development of artificial intelligence-based tools-particularly for free-text data extraction-will be crucial to enhance the accuracy and scalability of future RWD research.

Hanieh Abedian Kalkhoran, Tobias Martinot, Lydia H N Schonewille et al. · 0 citations
Review Open access Jul 2026

Large language model-based extraction of rheumatoid arthritis clinical disease activity index from clinical notes

Abstract Objectives We extracted a validated disease activity measure in rheumatoid arthritis (RA), the Clinical Disease Activity Index (CDAI), from a large tertiary academic medical center electronic health record (EHR) using an automated large language model (LLM)-based approach without requiring model pretraining. Materials and Methods The New York Presbyterian/Columbia University Medical Center Clinical Data Warehouse contains EHR data for over 4.5 million patients. RA patients were identified using International Classification of Disease-9 (ICD-9) and ICD-10 codes. Expert-curated CDAI keywords were extracted from unstructured notes using an automated natural language processing (NLP) pipeline leveraging GPT-4o API, a HIPAA-compliant, institutionally approved LLM platform. Performance was evaluated against expert chart review. Results Among 2756 RA patients with notes, 1038 (37.7%) were seropositive, 796 (28.9%) were seronegative, and 922 (33.4%) had unknown serostatus. Clinical Disease Activity Index and its components were extracted in 15.4% (160/1038) of seropositive patients indicating remission or low disease activity. Clinical Disease Activity Index documentation was more frequent among patients with multiple notes and among faculty, with high extraction accuracy (precision/recall/F1 = 0.97). Discussion This represents the first attempt to employ a zero-shot, ChatGPT-powered LLM platform to extract RA disease activity measures from real-world EHR data. Although a low prevalence of documentation was noted, important distinctions were observed when patients were subgrouped by serostatus, level of training, and number of visits. Conclusion An LLM-based pipeline accurately extracted CDAI from a single large academic EHR, revealing infrequent real-world documentation.

Reid Weisberg, Ruoqi Yang, Iram Kamdar et al. · 0 citations
Review Open access Jul 2026

AI-Assisted Clinical Data Abstraction From Electronic Health Records: Retrospective Concordance Study

Abstract Background Manual chart abstraction from electronic health records is a critical step in clinical outcomes research but is time-intensive and prone to human error. Advances in artificial intelligence (AI), particularly large language models, offer the potential to automate the extraction of structured data from unstructured clinical documentation with improved efficiency and consistency. Objective This study aimed to evaluate the accuracy and efficiency of an AI-assisted approach for extracting patient-reported outcomes from clinical notes compared with traditional human abstraction. Methods We conducted a retrospective study of 26 patients treated with low-dose radiation therapy for osteoarthritis. Human reviewers abstracted numeric rating scale (NRS; 0‐10) pain scores at baseline, the end of treatment, and the first follow-up, and von Pannewitz score (VPS; 0‐4) improvement scores at posttreatment time points. A HIPAA (Health Insurance Portability and Accountability Act)–compliant generative pretrained transformer–based AI system was prompted to extract the same end points from clinical notes. Concordance was assessed using exact match rates, the intraclass correlation coefficient for the NRS, and weighted Cohen κ for the VPS. The time required for AI vs manual abstraction was recorded. The AI system was not trained or fine-tuned on study data, and performance was evaluated directly against human abstraction to reflect real-world deployment. Results The AI system demonstrated high concordance with human abstraction, achieving an exact match rate of 92% for the NRS (95% CI 84‐96; intraclass correlation coefficient=0.96) and 94% for the VPS (95% CI 84‐98; κ=0.91). All discrepancies were minor, and no spurious values were generated. The AI system identified 1 clinically relevant data point missed during manual review. Average abstraction time per patient decreased from approximately 30 minutes to 2 minutes, representing time savings of >90%. The system also captured trends in analgesic use, but these results were not statistically significant, including reductions without escalation. Conclusions AI-assisted data abstraction demonstrated high concordance with human review in this single-institution cohort while substantially reducing the time requirements. These findings support the feasibility of AI-assisted abstraction workflows, although further validation across larger and more diverse datasets is needed.

Camille Sarah Schwartz, M. J. Anderson, K. Moakler et al. · 0 citations
Open access Jul 2026

Benchmarking large language models for clinical data extraction from Portuguese medical notes in a university hospital

Extracting structured data from electronic health records (EHRs) remains a major challenge, particularly in non-English and resource-constrained healthcare systems. This study benchmarks multiple large language models (LLMs) for the automated extraction of structured clinical variables from Portuguese-language medical notes under limited computational resources. We evaluated five LLMs (GPT-4o mini, DeepSeek-V3, Mixtral-8x7B, LLaMA 8B, and Qwen-32B) against a manually curated dataset of cardiology and infectiology outpatient records. Models were deployed in quantized versions to optimize computational efficiency. Outputs were compared with human annotations using F1 score, balanced accuracy, and recall. Among the tested models, Qwen-32B achieved the highest performance in both the infectiology domain (balanced accuracy = 0.91 [0.07]) and cardiology domain (balanced accuracy = 0.89 [0.07]). Performance varied by clinical variable, with better results for frequently and consistently documented conditions (e.g., diabetes) and lower accuracy for complex or infrequent variables (e.g., tumors). Extraction time ranged from 0.9 to 24.2 minutes per patient, depending on clinical domain and model. These findings support the feasibility of applying LLM-based natural language processing tools in resource-limited, non-English healthcare settings. Future research should assess emerging high-parameter models and explore additional clinical domains.

Breno Gabriel Araújo Sampaio de Jesus, Tomaz Castrillon Figueiredo, Clariele de Almeida Pereira et al. · 1 citation
Review Open access Aug 2026

Large Language Models in Adverse Drug Reaction Detection and Pharmacovigilance: A Systematic Review of Current Applications, Challenges, and Future Directions

Background/Objectives: Pharmacovigilance workflows rely heavily on unstructured text across diverse sources. Here, we systematically reviewed how large language models (LLMs) are being explored as support tools for adverse drug reaction (ADR) detection, extraction, triage, and documentation, highlighting their potential for precision medicine and big data-enabled safety monitoring. Methods: Following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses 2020 guidelines, we systematically searched PubMed, Scopus, and Web of Science for studies published between January 2022 and March 2026. Ultimately, 83 empirical studies satisfied the inclusion criteria. A narrative synthesis was conducted to address methodological heterogeneity across these studies. Results: LLM applications were concentrated in constrained information-extraction and classification tasks, including signal evaluation, clinical-note extraction, social media surveillance, and literature screening. Quantitative performance varied substantially by system design: error-correction prompting yielded an F1-score of 0.921 for ADR named entity recognition, whereas retrieval-augmented generation improved data-retrieval accuracy from 8.3% to 78.3%. Most studies were retrospective, benchmark-based, or proof-of-concept evaluations. Across 581 paired pre-consensus domain judgements, observed inter-rater agreement was 90.4% and Cohen’s κ was 0.837 (95% CI 0.772–0.895). Hallucination, low specificity, prompt sensitivity, narrow datasets, and weak external validation remained common limitations. Conclusions: Current evidence supports supervised, task-specific applications of LLMs for extraction, triage, retrieval, and documentation rather than autonomous pharmacovigilance decision-making. Prospective evaluation, external validation, transparent reporting, and accountable human oversight are required before high-stakes clinical or regulatory deployment.

Tae You Kim, Won-Sik Oh, Dong-Hwa Jeong · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 17, 2026

Q&A: Rethinking how innovation happens

In his latest book, Professor Eugene Fitzgerald examines the forces that turn breakthroughs into value — and why innovation resists simple formulas.