Skip to content
Review Open access

Performance of large language models on screening titles and abstracts of urology-related systematic reviews

Jul 2026 · McMaster University Medical Journal · Vol 22, pp. 52 · 0 citations

TL;DR

Locally deployed LLMs can screen titles and abstracts of urology-related systematic reviews with moderate accuracy, and next steps to improve LLM performance include creating random subsamples from each systematic review to perform calibration screening runs with the models.

Abstract

Background: Systematic reviews summarize research evidence to inform clinical practice guidelines and health policy, but the review process is labour-intensive. Authors often screen thousands of abstracts to identify relevant studies to include in their review. Large Language Models (LLMs) may improve title and abstract screening efficiency, but their performance compared to humans require evaluation. Objective: To assess whether locally deployed LLMs accurately screen titles and abstracts of urology-related systematic reviews. Methods: We used LLMs to screen titles and abstracts from three Cochrane reviews in urology totaling 8020 records. Records were screened using predefined criteria, as stated in the Cochrane manuscripts, using three strategies: Qwen3 30B A3B Instruct 2507 alone, Ministral 3 14B Instruct 2512 alone, and their ensemble. A “first-ahead-by-k” approach (k = 2) was used with two models in parallel. We compared their outputs, and once a classification led by a margin of K (i.e., it appeared two more times than any alternative in the running tally) we selected that label as the consensus. Strategy performance metrics were calculated against the human-defined gold standard. Results: Specificity consistently exceeded 90% across strategies (Figure 1). Pooled estimates of sensitivity across the three reviews for strategies were: Ministral 3 14B, 71.1% (95% CI, 64.9-76.6); Qwen3 30B A3B, 53.1% (46.6-59.4); and their ensemble, 72.2% (66.1-77.7). Conclusion: Locally deployed LLMs can screen titles and abstracts with moderate accuracy. Next steps to improve LLM performance include creating random subsamples from each systematic review to perform calibration screening runs with the models, reviewing discrepancies between the model and human-defined gold standard to determine why the performance was variable, and retrieving the exact title and abstract screening criteria used by the authors.

Read PDF

Similar papers

Review Aug 2026

ENTGPT: Applying Large Language Models to Systematic Review Screening With the Novel STARR Protocol.

Performance of ENTGPT accurately replicated human reviewers in article selection and data extraction for an otolaryngology SLR using the STARR and traditional protocols suggests that LLMs could be employed to significantly streamline the SLR process, potentially saving substantial time and resources for researchers.

Akash Kapoor, Ben Baranker, I. Alter et al. · 0 citations
Review Open access Sep 2026

USING LARGE LANGUAGE MODELS FOR LITERATURE SEARCH IN CARDIOVASCULAR SURGERY SYSTEMATIC REVIEWS AND META-ANALYSES

Highlights Large Language Models (LLMs) enable automation of initial abstract screening in systematic reviews, significantly reducing manual workload. The effectiveness of LLMs heavily depends on prompt engineering, which must clearly translate inclusion and exclusion criteria into actionable instructio...

Shatskiy Alexander S., Ehab M. Deigheidy, S. E. Masyutina et al. · 0 citations
#small language model Review Open access Aug 2026

Toward Automating the Selection of Articles Reporting EQ-5D Data for Systematic Literature Reviews Using Large Language Models: Algorithm Development and Evaluation Study

The models reproduced human screening tendencies despite the small dataset size, demonstrating the technical feasibility of LLM-assisted article selection and providing the first demonstration of LLM-assisted identification of EQ-5D data in biomedical literature.

Gábor Kertész, J. Czere, Z. Zrubka et al. · 0 citations
Review Open access Aug 2026

Performance of Two AI Approaches in ASReview Compared With Manual Screening for Dementia Care Literature Screening: Comparative Analysis

ASReview can support workload reduction in title and abstract screening, but the evaluated ASReview approaches did not retrieve all final included studies from the original dementia care scoping review, suggesting that the evaluated ASReview configurations may be insufficient for reviews in which near-complete retrieva...

Dirk Steijger, Stella Thissen, Sil Aarts et al. · 0 citations
#small language model Review Open access Sep 2026

A large language model for risk-of-bias assessment in systematic reviews of prognosis studies in clinical neurology

With targeted methodological refinements - including standardization of QUIPS implementation and validation against expert ratings - automated ROB assessments may meaningfully reduce time and cost of systematic reviews of prognosis studies in neurology and beyond.

S. Kampman, K. Braun, W. Otte · 0 citations
Review Open access Oct 2026

Title and abstract screening for systematic reviews with Jev, a System One model: comparison with generative large language models

Large language models (LLMs) screen titles and abstracts without review-specific training, but generating screening decisions as text takes processing time and incurs API charges. We evaluated Jev, a non-generative model returning classification probabilities, on 4527 records from two systematic reviews of bipolar diso...

K. Matsui, Y. Takaesu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.