Skip to content
Review

ENTGPT: Applying Large Language Models to Systematic Review Screening With the Novel STARR Protocol.

Aug 2026 · The Laryngoscope · 0 citations · 14 references
Medicine

TL;DR

Performance of ENTGPT accurately replicated human reviewers in article selection and data extraction for an otolaryngology SLR using the STARR and traditional protocols suggests that LLMs could be employed to significantly streamline the SLR process, potentially saving substantial time and resources for researchers.

Abstract

Objectives

Systematic literature reviews (SLRs) are time-intensive and resource-consuming. While large language models (LLMs) have shown promise in encoding clinical knowledge, evidence for their performance on complex text analysis, necessary for SLRs in otolaryngology, remains limited. In this proof-of-concept study, we investigate an LLM's performance in screening relevant articles for a SLR with the novel screening of title and abstracts, reevaluation, and full-text review (STARR) protocol.

Methods

ENTGPT (based on GPT-4o) was compared to two human reviewers in article inclusion/exclusion decisions using the traditional and STARR screening protocols. ENTGPT was provided with inclusion and exclusion criteria, titles, abstracts, and full texts (if available) of the 850 articles retrieved in the original search. The model's decisions were compared to those made by two human reviewers.

Results

ENTGPT, using the STARR protocol, achieved 99.87% accuracy in article classification compared with human reviewers (95% CI: 0.99-1.0), including 100% specificity and 95% sensitivity. When using the traditional protocol, sensitivity declined to 35%. ENTGPT, using the traditional protocol, achieved 99.47% accuracy in article classification compared with human reviewers (95% CI: 0.99-1.0), including 100% specificity and 35% sensitivity. When using the STARR protocol, accuracy improved to 99.87% and sensitivity markedly increased to 95%.

Conclusions

ENTGPT accurately replicated human reviewers in article selection and data extraction for an otolaryngology SLR using the STARR and traditional protocols. This performance suggests that LLMs could be employed to significantly streamline the SLR process, potentially saving substantial time and resources for researchers. LEVEL OF EVIDENCE N/A.

View source

Similar papers

Review Open access Sep 2026

USING LARGE LANGUAGE MODELS FOR LITERATURE SEARCH IN CARDIOVASCULAR SURGERY SYSTEMATIC REVIEWS AND META-ANALYSES

Highlights Large Language Models (LLMs) enable automation of initial abstract screening in systematic reviews, significantly reducing manual workload. The effectiveness of LLMs heavily depends on prompt engineering, which must clearly translate inclusion and exclusion criteria into actionable instructio...

Shatskiy Alexander S., Ehab M. Deigheidy, S. E. Masyutina et al. · 0 citations
#small language model Review Open access Sep 2026

A large language model for risk-of-bias assessment in systematic reviews of prognosis studies in clinical neurology

With targeted methodological refinements - including standardization of QUIPS implementation and validation against expert ratings - automated ROB assessments may meaningfully reduce time and cost of systematic reviews of prognosis studies in neurology and beyond.

S. Kampman, K. Braun, W. Otte · 0 citations
Review Open access Sep 2026

Large Language Model versus Clinician Written Summaries of Research Papers.

An enterprise LLM, prompted in POEM style, produced accurate, low-error clinical summaries that matched or exceeded expert-edited POEMs and were generally preferred by reviewers, though further research is needed to assess broader applicability and impact.

Richard Guthmann, Robert Martin, Erin Lee et al. · 0 citations
Review Open access Sep 2026

Large language models for patient-facing pathology report interpretation: A scoping review.

LLM-based patient-facing pathology report interpretation shows potential to bridge specialist pathology language and patient communication, and evaluation should extend beyond readability to include fidelity to the original pathology report, patient understanding, safety, and usability.

Chen Wang, Jie Hao, Si-Jia Zhang et al. · 0 citations
#large language models Review Sep 2026

Large language models for literature screening in conceptually complex and interdisciplinary reviews

It is suggested that current LLMs can support literature screening, but their reliability depends strongly on the conceptual clarity of the review task and the structure of eligibility criteria, as well as practical implications for the transparent and responsible use of LLMs in informetric research and systematic revi...

Ni Cheng, Heng Dong, Xuan Han et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.