Skip to content
Review Open access

Performance of Two AI Approaches in ASReview Compared With Manual Screening for Dementia Care Literature Screening: Comparative Analysis

Aug 2026 · JMIR Formative Research · Vol 10 · 0 citations · 46 references
Medicine

TL;DR

ASReview can support workload reduction in title and abstract screening, but the evaluated ASReview approaches did not retrieve all final included studies from the original dementia care scoping review, suggesting that the evaluated ASReview configurations may be insufficient for reviews in which near-complete retrieval of relevant evidence is required.

Abstract

Abstract Background Literature reviews rely on rigorous title and abstract screening by researchers, which is time-consuming. AI-assisted literature screening tools have been proposed to improve efficiency by prioritizing titles and abstracts with the highest likelihood of meeting the inclusion criteria, thereby reducing the need to screen all records. Objective This study aims to evaluate the performance of two AI-assisted screening approaches in ASReview (version 1.3; Department of Methodology and Statistics, Utrecht University) compared with manual title and abstract screening in a previously completed and published scoping review on how AI can support the quality of life in people with dementia. Methods This study used a dataset of 4690 titles and abstracts from a published scoping review. The manual screening decisions of the scoping review served as the reference standard. Both ASReview approaches were applied by the same author who conducted the majority of the original manual title and abstract screening. Approach A used a simpler model with minimal prior input, whereas approach B used a more advanced model with a larger training set. Both ASReview approaches applied predefined stopping rules: (1) more than 10% of the dataset to be screened; and (2) 50 consecutive irrelevant titles and abstracts. Performance was evaluated in terms of sensitivity, specificity, precision, accuracy, and screening time. 95% CIs were calculated for sensitivity, specificity, precision, and accuracy. Agreement between manual and ASReview approaches was assessed using the Cohen κ, and differences in how manual and both ASReview approaches classified titles and abstracts were examined using the McNemar test. Performance and agreement were calculated at two levels: (1) after title and abstract screening and (2) after full-text inclusion. Results Manual screening identified 283 titles and abstracts for full-text review and resulted in 30 final included studies, requiring 19 hours. Of the 4690 titles and abstracts, approach A screened 830 (17.7%) in 4.3 hours and retrieved 16 of the 30 (sensitivity 0.53, 95% CI 0.36‐0.70) final included studies, whereas approach B screened 798 (17.0%) in 5.5 hours and retrieved 21 of the 30 (sensitivity 0.70, 95% CI 0.52‐0.83) final included studies. Although both ASReview approaches showed high specificity and accuracy, these metrics should be interpreted cautiously because the dataset was highly imbalanced and contained relatively few relevant titles and abstracts. McNemar tests showed significant directional imbalance (P<.001): ASReview missed more manually selected titles and abstracts at level 1, whereas at level 2, ASReview more often labeled titles and abstracts not included in the final review as relevant. Conclusions ASReview can support workload reduction in title and abstract screening, but the evaluated ASReview approaches did not retrieve all final included studies from the original dementia care scoping review. These findings suggest that the evaluated ASReview configurations may be insufficient for reviews in which near-complete retrieval of relevant evidence is required.

Read PDF

Similar papers

Review

Journal of Health Economics and Outcomes Research

An AI-assisted SLR was conducted that mirrored a traditional SLR performed by humans only and assessed the performance of the AI tools employed, finding uneven results across stages reinforce the need for human oversight and clear validation methods.

Jing Wang-Silvanto, Mansee Jajoo, Rishi Ohri et al. · 0 citations
Review Open access Aug 2026

NEUROLOGICAL BIOMARKERS FOR EARLY DIAGNOSIS OF ALZHEIMER'S DISEASE: AN INTEGRATIVE LITERATURE REVIEW

Objective: To analyze the available scientific evidence on neurological biomarkers used for early diagnosis of Alzheimer's disease (AD), evaluating their diagnostic accuracy and predictive capacity for disease progression. Methods: This is an integrative literature review conducted according to the guidelines proposed...

João Pedro Gomes do Nascimento, Maria Luiza de Freitas Serafim, Rhayssa Mylla Barbosa Celestino et al. · 0 citations
Review Aug 2026

ENTGPT: Applying Large Language Models to Systematic Review Screening With the Novel STARR Protocol.

Performance of ENTGPT accurately replicated human reviewers in article selection and data extraction for an otolaryngology SLR using the STARR and traditional protocols suggests that LLMs could be employed to significantly streamline the SLR process, potentially saving substantial time and resources for researchers.

Akash Kapoor, Ben Baranker, I. Alter et al. · 0 citations
#small language model Review Open access Sep 2026

A large language model for risk-of-bias assessment in systematic reviews of prognosis studies in clinical neurology

With targeted methodological refinements - including standardization of QUIPS implementation and validation against expert ratings - automated ROB assessments may meaningfully reduce time and cost of systematic reviews of prognosis studies in neurology and beyond.

S. Kampman, K. Braun, W. Otte · 0 citations
#small language model Review Open access Aug 2026

Toward Automating the Selection of Articles Reporting EQ-5D Data for Systematic Literature Reviews Using Large Language Models: Algorithm Development and Evaluation Study

The models reproduced human screening tendencies despite the small dataset size, demonstrating the technical feasibility of LLM-assisted article selection and providing the first demonstration of LLM-assisted identification of EQ-5D data in biomedical literature.

Gábor Kertész, J. Czere, Z. Zrubka et al. · 0 citations
Review Open access Sep 2026

USING LARGE LANGUAGE MODELS FOR LITERATURE SEARCH IN CARDIOVASCULAR SURGERY SYSTEMATIC REVIEWS AND META-ANALYSES

Highlights Large Language Models (LLMs) enable automation of initial abstract screening in systematic reviews, significantly reducing manual workload. The effectiveness of LLMs heavily depends on prompt engineering, which must clearly translate inclusion and exclusion criteria into actionable instructio...

Shatskiy Alexander S., Ehab M. Deigheidy, S. E. Masyutina et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.