Jul 2026· Journal of Biomedical Informatics· pp.
105086
· 0 citations· 24 references
MedicineComputer Science
Abstract
Objective
Data extraction is among the most resource-intensive and error-prone stages of systematic review production. Large language models (LLMs) offer potential for automating or semi-automating this process, yet their performance characteristics remain incompletely characterised. This systematic review aimed to comprehensively evaluate LLM accuracy, reliability, and efficiency for data extraction in evidence synthesis, and to identify optimal implementation strategies.
Methods
We searched PubMed, Embase, Web of Science, and preprint servers (medRxiv, arXiv) through December 2025. Studies were eligible if they evaluated one or more LLMs for data extraction against a human reference standard and reported quantitative performance metrics. Two reviewers independently extracted data and assessed methodological quality using PROBAST + AI and reporting completeness using TRIPOD-LLM. Narrative synthesis was performed due to substantial heterogeneity precluding meta-analysis.
Results
Twenty-seven studies met inclusion criteria, evaluating models including GPT-4/4o (n = 15), Claude versions 2-3.5 (n = 10), Gemini (n = 3), and open-source alternatives including Llama, Mistral, Qwen, and DeepSeek. Overall accuracy ranged from 47% to 99.9%, with substantial heterogeneity by task type and data granularity. Categorical and string variables were extracted more reliably (74-96%) than numerical data (47-88%). Claude 3.5 Sonnet achieved high accuracy in an assistive workflow (91.0%; 95% CI: 90.4-91.6%), exceeding human-only extraction (89.0%). Claude models outperformed GPT in head-to-head comparisons (OR 1.70 for event counts). Omissions were the dominant error type (60-74%), with hallucination rates of only 0.08-6%, challenging widespread fabrication concerns. Time savings of 33% to 87% were reported, although most included studies did not quantitatively assess efficiency. Methodological quality was generally robust, with 74.1% of studies rated low risk of bias under PROBAST + AI. Mean TRIPOD-LLM compliance was 88.5%, though gaps in inference settings and model version documentation were common.
Conclusion
LLMs demonstrate promising but variable performance for data extraction in evidence synthesis. Current evidence supports their integration as assistive tools within dual-extraction workflows requiring human verification, rather than as autonomous extractors. Categorical data is extracted more reliably than numerical outcomes, and few-shot prompting with structured output formats consistently improves performance. Standardised benchmarks and prospective comparative studies remain priorities for future research.
A suite of automated tools for automated batch processing that provide decision rationales and evidence enhances transparency and allows for human verification of AI decisions and provides a suite of automated tools for key SR tasks.
Yi-Ran Liu, Xi-Ling Wang, Zi-Xuan Zhou et al.· Journal of Evaluation In Cli...· 0 citations
Current evidence supports supervised, task-specific applications of LLMs for extraction, triage, retrieval, and documentation rather than autonomous pharmacovigilance decision-making, highlighting their potential for precision medicine and big data-enabled safety monitoring.
Tae You Kim, Won-Sik Oh, Dong-Hwa Jeong· Diagnostics· 0 citations
The models reproduced human screening tendencies despite the small dataset size, demonstrating the technical feasibility of LLM-assisted article selection and providing the first demonstration of LLM-assisted identification of EQ-5D data in biomedical literature.
Gábor Kertész, J. Czere, Z. Zrubka et al.· JMIR Formative Research· 0 citations
SciLitBench identifies a practical boundary between high-recall screening and evidence-complete extraction and provides a reproducible resource for evaluating LLM-assisted evidence synthesis.
It is demonstrated that modern LLMs with reasoning capabilities can effectively support real-time clinical evidence synthesis without task-specific fine-tuning, offering a pathway toward scalable automated systems for clinical trial interpretation that could substantially reduce the evidence-to-practice gap in medical...
Shobanapriyan Chandrasegaran, Amal Htait· International Conference on...· 0 citations
An open-source, end-to-end pipeline is built to simulate LLM performance in a hypothetical scenario where they were available to inform COVID-19 models developed during the first four months of 2020, and shows that current models can extract transmission parameters from unstructured literature accurately enough to info...
X. Yang, C. Y. Lee, B. Quilty et al.· medRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.