GenAI can improve efficiency across multiple SLR tasks when used in hybrid human-AI workflows when used in hybrid human-AI workflows, and current evidence supports targeted, task-specific adoption with transparent reporting and human oversight, rather than full automation.
Abstract
Objectives
Systematic literature reviews (SLRs) underpin life sciences research but are resource intensive. Generative artificial intelligence, particularly large language models (LLMs), may accelerate key SLR tasks, yet performance and reliability for evidence synthesis remain unclear. This manuscript aims to review current evidence on GenAI performance across core SLR tasks.
Methods
We conducted a PRISMA-adapted rapid evidence assessment of English-language biomedical studies published from November 2022 to July 2025 evaluating GenAI or LLMs for systematic literature review tasks, including search strategy development, title/abstract screening, full-text screening, data extraction, risk-of-bias assessment, qualitative synthesis, report writing, and end-to-end review generation. Findings were summarized qualitatively by task.
Results
Among 115 included studies, evidence supporting the use of GenAI was strongest for title/abstract screening (n=51) and data extraction (n=33). Selected high-quality evaluations reported sensitivities ≥90%, workload reductions of 27-71%, and human-comparable or superior performance in calibrated human-in-the-loop workflows. Evidence for full-text screening (n=15) and risk-of-bias assessment (n=17) was more variable, showing gains in structured or fine-tuned implementations but persistent limitations in specificity and nuanced judgment. For search strategy development, qualitative synthesis, and report writing, GenAI was most effective as a supportive tool; fully autonomous end-to-end SLR generation was unreliable.
Conclusions
GenAI can improve efficiency across multiple SLR tasks when used in hybrid human-AI workflows. Current evidence supports targeted, task-specific adoption with transparent reporting and human oversight, rather than full automation.
OBJECTIVE
Data extraction is among the most resource-intensive and error-prone stages of systematic review production. Large language models (LLMs) offer potential for automating or semi-automating this process, yet their performance characteristics remain incompletely characterised. This systematic review aimed to com...
Ravi Shankar, Amaevia Lim, Xu Qian· Journal of Biomedical Inform...· 0 citations
Artificial intelligence-driven tools performed well for title/abstract screening and general data extraction but were less accurate for full-text screening and interpretation of modeling choices.
L. Bloudek, Allie Cichewicz, K. Patel et al.· PharmacoEconomics (Auckland)· 0 citations
It was revealed that most GenAI applications in healthcare rely on general purpose LLMs to provide treatment recommendations, and future research should prioritize the development of interpretable, domain-specific models and rigorous clinical trials to ensure safe and effective integration into healthcare settings.
Leonides Medeiros, Maicon Herverton Lino Ferreira da Silva Barros, K. H. de Carvalho Monteiro et al.· Journal of Healthcare Inform...· 0 citations
A suite of automated tools for automated batch processing that provide decision rationales and evidence enhances transparency and allows for human verification of AI decisions and provides a suite of automated tools for key SR tasks.
Yi-Ran Liu, Xi-Ling Wang, Zi-Xuan Zhou et al.· Journal of Evaluation In Cli...· 0 citations
SciLitBench identifies a practical boundary between high-recall screening and evidence-complete extraction and provides a reproducible resource for evaluating LLM-assisted evidence synthesis.
It is demonstrated that modern LLMs with reasoning capabilities can effectively support real-time clinical evidence synthesis without task-specific fine-tuning, offering a pathway toward scalable automated systems for clinical trial interpretation that could substantially reduce the evidence-to-practice gap in medical...
Shobanapriyan Chandrasegaran, Amal Htait· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.