Skip to content
Review

Performance of large language models in data extraction for evidence synthesis: A systematic review

Jul 2026 · Journal of Biomedical Informatics · pp. 105086 · 0 citations · 24 references
Medicine Computer Science

Abstract

Objective

Data extraction is among the most resource-intensive and error-prone stages of systematic review production. Large language models (LLMs) offer potential for automating or semi-automating this process, yet their performance characteristics remain incompletely characterised. This systematic review aimed to comprehensively evaluate LLM accuracy, reliability, and efficiency for data extraction in evidence synthesis, and to identify optimal implementation strategies.

Methods

We searched PubMed, Embase, Web of Science, and preprint servers (medRxiv, arXiv) through December 2025. Studies were eligible if they evaluated one or more LLMs for data extraction against a human reference standard and reported quantitative performance metrics. Two reviewers independently extracted data and assessed methodological quality using PROBAST + AI and reporting completeness using TRIPOD-LLM. Narrative synthesis was performed due to substantial heterogeneity precluding meta-analysis.

Results

Twenty-seven studies met inclusion criteria, evaluating models including GPT-4/4o (n = 15), Claude versions 2-3.5 (n = 10), Gemini (n = 3), and open-source alternatives including Llama, Mistral, Qwen, and DeepSeek. Overall accuracy ranged from 47% to 99.9%, with substantial heterogeneity by task type and data granularity. Categorical and string variables were extracted more reliably (74-96%) than numerical data (47-88%). Claude 3.5 Sonnet achieved high accuracy in an assistive workflow (91.0%; 95% CI: 90.4-91.6%), exceeding human-only extraction (89.0%). Claude models outperformed GPT in head-to-head comparisons (OR 1.70 for event counts). Omissions were the dominant error type (60-74%), with hallucination rates of only 0.08-6%, challenging widespread fabrication concerns. Time savings of 33% to 87% were reported, although most included studies did not quantitatively assess efficiency. Methodological quality was generally robust, with 74.1% of studies rated low risk of bias under PROBAST + AI. Mean TRIPOD-LLM compliance was 88.5%, though gaps in inference settings and model version documentation were common.

Conclusion

LLMs demonstrate promising but variable performance for data extraction in evidence synthesis. Current evidence supports their integration as assistive tools within dual-extraction workflows requiring human verification, rather than as autonomous extractors. Categorical data is extracted more reliably than numerical outcomes, and few-shot prompting with structured output formats consistently improves performance. Standardised benchmarks and prospective comparative studies remain priorities for future research.

View source

Similar papers

Review Open access Sep 2026

Performance and Consistency of Large Language Models in Key Labor-Intensive Tasks of Systematic Reviews.

A suite of automated tools for automated batch processing that provide decision rationales and evidence enhances transparency and allows for human verification of AI decisions and provides a suite of automated tools for key SR tasks.

Yi-Ran Liu, Xi-Ling Wang, Zi-Xuan Zhou et al. · 0 citations
Review Open access Aug 2026

Large Language Models in Adverse Drug Reaction Detection and Pharmacovigilance: A Systematic Review of Current Applications, Challenges, and Future Directions

Current evidence supports supervised, task-specific applications of LLMs for extraction, triage, retrieval, and documentation rather than autonomous pharmacovigilance decision-making, highlighting their potential for precision medicine and big data-enabled safety monitoring.

Tae You Kim, Won-Sik Oh, Dong-Hwa Jeong · 0 citations
#small language model Review Open access Aug 2026

Toward Automating the Selection of Articles Reporting EQ-5D Data for Systematic Literature Reviews Using Large Language Models: Algorithm Development and Evaluation Study

The models reproduced human screening tendencies despite the small dataset size, demonstrating the technical feasibility of LLM-assisted article selection and providing the first demonstration of LLM-assisted identification of EQ-5D data in biomedical literature.

Gábor Kertész, J. Czere, Z. Zrubka et al. · 0 citations
Review 2026

Enhancing Clinical Trial Analysis through Large Language Models for Multi-Evidence Natural Language Inference

It is demonstrated that modern LLMs with reasoning capabilities can effectively support real-time clinical evidence synthesis without task-specific fine-tuning, offering a pathway toward scalable automated systems for clinical trial interpretation that could substantially reduce the evidence-to-practice gap in medical...

Shobanapriyan Chandrasegaran, Amal Htait · 0 citations
Review Open access Sep 2026

Using large language models to facilitate literature review and data extraction for infectious disease models: COVID-19 as a test case

An open-source, end-to-end pipeline is built to simulate LLM performance in a hypothetical scenario where they were available to inform COVID-19 models developed during the first four months of 2020, and shows that current models can extract transmission parameters from unstructured literature accurately enough to info...

X. Yang, C. Y. Lee, B. Quilty et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.