Skip to content

Large language models for literature screening in conceptually complex and interdisciplinary reviews

Sep 2026 · Aslib Journal of Information Management · pp. 1-23 · 0 citations · 58 references
Meta-analysis and systematic reviews

TL;DR

It is suggested that current LLMs can support literature screening, but their reliability depends strongly on the conceptual clarity of the review task and the structure of eligibility criteria, as well as practical implications for the transparent and responsible use of LLMs in informetric research and systematic review practice.

Abstract

This study examines the capabilities and limitations of large language models (LLMs) for literature screening in systematic reviews involving conceptually diffuse and interdisciplinary topics. Two screening datasets were constructed. Four LLMs, Claude 4.6, DeepSeek V4 Pro, Gemini 3.1 Pro, and GPT-5.5, were evaluated using a staged, rule-guided prompting workflow covering title screening, abstract screening, and full-text assessment. Model performance was assessed using accuracy, precision, recall, specificity, the F1-score, and Cohen's kappa, supplemented by stage-wise screening flow analysis and qualitative discrepancy analysis between model and human reviewer decisions. Model performance varied substantially across datasets and models. In the conceptually diffuse Dataset A, all models achieved high specificity, but positive-class performance was more limited. DeepSeek V4 Pro achieved the highest accuracy, precision, F1-score, and Cohen's kappa, whereas Claude 4.6 achieved the highest recall. In the more clearly bounded Dataset B, recall was higher and model behavior was more convergent, with GPT-5.5 showing the best overall balance across performance indicators. Stage-wise analysis showed that models differed in where they made inclusion and exclusion decisions, while discrepancy analysis indicated that errors were mainly related to conceptual scope confusion, context misidentification, and underestimation of analytical depth. These findings suggest that current LLMs can support literature screening, but their reliability depends strongly on the conceptual clarity of the review task and the structure of eligibility criteria. This study extends the literature on LLM-assisted screening by focusing on conceptually abstract and interdisciplinary review tasks. This study introduces a staged, Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA)-aligned screening workflow and conducts an error analysis that explains how LLM screening can fail in semantically diffuse settings. The study provides practical implications for the transparent and responsible use of LLMs in informetric research and systematic review practice.

View source

Similar papers

Review Open access Sep 2026

Performance and Consistency of Large Language Models in Key Labor-Intensive Tasks of Systematic Reviews.

A suite of automated tools for automated batch processing that provide decision rationales and evidence enhances transparency and allows for human verification of AI decisions and provides a suite of automated tools for key SR tasks.

Yi-Ran Liu, Xi-Ling Wang, Zi-Xuan Zhou et al. · 0 citations
Review Open access 2026

Large Language Model-Based Automated Assessment: A Systematic Review, Taxonomy, and Implications for Personalized Learning

Research on large language model (LLM)-based automated assessment (AA) has expanded rapidly. Nevertheless, the literature remains fragmented across contributions, models, implementation configurations, datasets, and evaluation metrics, complicating efforts to identify approaches suitable for personalized learning. This...

H. D. Septama, A. E. Permanasari, R. Ferdiana · 0 citations
Review Open access Sep 2026

USING LARGE LANGUAGE MODELS FOR LITERATURE SEARCH IN CARDIOVASCULAR SURGERY SYSTEMATIC REVIEWS AND META-ANALYSES

Highlights Large Language Models (LLMs) enable automation of initial abstract screening in systematic reviews, significantly reducing manual workload. The effectiveness of LLMs heavily depends on prompt engineering, which must clearly translate inclusion and exclusion criteria into actionable instructio...

Shatskiy Alexander S., Ehab M. Deigheidy, S. E. Masyutina et al. · 0 citations
Review Aug 2026

Evaluating human and LLM screening workflows in a conceptually complex scoping review: Recall--workload trade-offs and run-to-run consistency

Compared human and LLM title-and-abstract screening workflows in a preregistered study embedded in a conceptually complex scoping review, finding that large language models are better suited to validated, auditable, human-supervised workflows than autonomous exclusion.

Nikol Figalová, Lynn Huestegge, Anne Böckler-Raettig · 0 citations
Review Open access Oct 2026

Title and abstract screening for systematic reviews with Jev, a System One model: comparison with generative large language models

Large language models (LLMs) screen titles and abstracts without review-specific training, but generating screening decisions as text takes processing time and incurs API charges. We evaluated Jev, a non-generative model returning classification probabilities, on 4527 records from two systematic reviews of bipolar diso...

K. Matsui, Y. Takaesu · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.