A complete workflow that can be adopted for new, unlabelled reviews, using open-source LLMs small enough to run on a high-end consumer laptop, and provided as an open-source R package is offered.
Abstract
Large language models (LLMs) can ease the work of screening titles and abstracts for systematic reviews, but obtaining reliable results requires researchers to make practical choices about which LLMs to use, how to combine their scores into a ranking, and how far down that ranking to read. We aimed to identify a general-purpose workflow that screens accurately, minimises human review effort, and generalises across environmental literature corpora. We ran an ensemble of five open-source LLMs across ten human-annotated systematic reviews from the field of ecology and environmental science spanning 19,155 studies. We then asked: (1) how well an ensemble of LLMs ranks relevant papers above irrelevant ones, and (2) where a human reviewer should stop working down that ranked list. A four-LLM ensemble chosen without any labels came close, on every review, to the best ranking achievable with that review’s annotations (mean Average Precision 0.66 versus 0.68). We tested different rules for when to stop human review, finding that a single SAFE stopping rule chosen in advance recovered ≥ 95% of relevant records on all ten reviews while requiring a human to screen 56% of the corpus on average. Against the established open-source active-learning tool ASReview, this label-free ranking reached the same recall target at lower human workload on seven of the ten reviews (39% versus 42% of the corpus on average). The paper offers a complete workflow that can be adopted for new, unlabelled reviews, using open-source LLMs small enough to run on a high-end consumer laptop, and we provide it as an open-source R package.
SciLitBench identifies a practical boundary between high-recall screening and evidence-complete extraction and provides a reproducible resource for evaluating LLM-assisted evidence synthesis.
Research on large language model (LLM)-based automated assessment (AA) has expanded rapidly. Nevertheless, the literature remains fragmented across contributions, models, implementation configurations, datasets, and evaluation metrics, complicating efforts to identify approaches suitable for personalized learning. This...
H. D. Septama, A. E. Permanasari, R. Ferdiana· IEEE Access· 0 citations
An evaluation framework that accounts for class imbalance is proposed, i.e., the natural prevalence of excluded articles relative to included articles in SRs, and PromptSR, a tool designed to support prompt experimentation, experiment management, and result analysis for LLM-based screening are introduced.
G.Aravind Kumar, Luciano Marchezan, G. Genois et al.· 0 citations
A suite of automated tools for automated batch processing that provide decision rationales and evidence enhances transparency and allows for human verification of AI decisions and provides a suite of automated tools for key SR tasks.
Yi-Ran Liu, Xi-Ling Wang, Zi-Xuan Zhou et al.· Journal of Evaluation In Cli...· 0 citations
It is suggested that current LLMs can support literature screening, but their reliability depends strongly on the conceptual clarity of the review task and the structure of eligibility criteria, as well as practical implications for the transparent and responsible use of LLMs in informetric research and systematic revi...
Ni Cheng, Heng Dong, Xuan Han et al.· Aslib Journal of Information...· 0 citations
AI and Large Language Models (LLMs) are increasingly applied to modern information processing, and ecology and biodiversity are no exception. The OneStop (2025) project seeks to minimise the introduction, establishment and spread of terrestrial invasive alien species. As part of this project, we are investigating how L...
Matthew Coole, Tom August· ARPHA Conference Abstracts· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.