Jul 2026· McMaster University Medical Journal· Vol 22, pp. 52· 0 citations
TL;DR
Locally deployed LLMs can screen titles and abstracts of urology-related systematic reviews with moderate accuracy, and next steps to improve LLM performance include creating random subsamples from each systematic review to perform calibration screening runs with the models.
Abstract
Background: Systematic reviews summarize research evidence to inform clinical practice guidelines and health policy, but the review process is labour-intensive. Authors often screen thousands of abstracts to identify relevant studies to include in their review. Large Language Models (LLMs) may improve title and abstract screening efficiency, but their performance compared to humans require evaluation.
Objective: To assess whether locally deployed LLMs accurately screen titles and abstracts of urology-related systematic reviews.
Methods: We used LLMs to screen titles and abstracts from three Cochrane reviews in urology totaling 8020 records. Records were screened using predefined criteria, as stated in the Cochrane manuscripts, using three strategies: Qwen3 30B A3B Instruct 2507 alone, Ministral 3 14B Instruct 2512 alone, and their ensemble. A “first-ahead-by-k” approach (k = 2) was used with two models in parallel. We compared their outputs, and once a classification led by a margin of K (i.e., it appeared two more times than any alternative in the running tally) we selected that label as the consensus. Strategy performance metrics were calculated against the human-defined gold standard.
Results: Specificity consistently exceeded 90% across strategies (Figure 1). Pooled estimates of sensitivity across the three reviews for strategies were: Ministral 3 14B, 71.1% (95% CI, 64.9-76.6); Qwen3 30B A3B, 53.1% (46.6-59.4); and their ensemble, 72.2% (66.1-77.7).
Conclusion: Locally deployed LLMs can screen titles and abstracts with moderate accuracy. Next steps to improve LLM performance include creating random subsamples from each systematic review to perform calibration screening runs with the models, reviewing discrepancies between the model and human-defined gold standard to determine why the performance was variable, and retrieving the exact title and abstract screening criteria used by the authors.
Performance of ENTGPT accurately replicated human reviewers in article selection and data extraction for an otolaryngology SLR using the STARR and traditional protocols suggests that LLMs could be employed to significantly streamline the SLR process, potentially saving substantial time and resources for researchers.
Akash Kapoor, Ben Baranker, I. Alter et al.· The Laryngoscope· 0 citations
Highlights
Large Language Models (LLMs) enable automation of initial abstract screening in systematic reviews, significantly reducing manual workload.
The effectiveness of LLMs heavily depends on prompt engineering, which must clearly translate inclusion and exclusion criteria into actionable instructio...
Shatskiy Alexander S., Ehab M. Deigheidy, S. E. Masyutina et al.· Complex Issues of Cardiovasc...· 0 citations
The models reproduced human screening tendencies despite the small dataset size, demonstrating the technical feasibility of LLM-assisted article selection and providing the first demonstration of LLM-assisted identification of EQ-5D data in biomedical literature.
Gábor Kertész, J. Czere, Z. Zrubka et al.· JMIR Formative Research· 0 citations
ASReview can support workload reduction in title and abstract screening, but the evaluated ASReview approaches did not retrieve all final included studies from the original dementia care scoping review, suggesting that the evaluated ASReview configurations may be insufficient for reviews in which near-complete retrieva...
Dirk Steijger, Stella Thissen, Sil Aarts et al.· JMIR Formative Research· 0 citations
With targeted methodological refinements - including standardization of QUIPS implementation and validation against expert ratings - automated ROB assessments may meaningfully reduce time and cost of systematic reviews of prognosis studies in neurology and beyond.
S. Kampman, K. Braun, W. Otte· medRxiv· 0 citations
Large language models (LLMs) screen titles and abstracts without review-specific training, but generating screening decisions as text takes processing time and incurs API charges. We evaluated Jev, a non-generative model returning classification probabilities, on 4527 records from two systematic reviews of bipolar diso...
K. Matsui, Y. Takaesu· medRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.