Aug 2026· The Laryngoscope· 0 citations· 14 references
Medicine
TL;DR
Performance of ENTGPT accurately replicated human reviewers in article selection and data extraction for an otolaryngology SLR using the STARR and traditional protocols suggests that LLMs could be employed to significantly streamline the SLR process, potentially saving substantial time and resources for researchers.
Abstract
Objectives
Systematic literature reviews (SLRs) are time-intensive and resource-consuming. While large language models (LLMs) have shown promise in encoding clinical knowledge, evidence for their performance on complex text analysis, necessary for SLRs in otolaryngology, remains limited. In this proof-of-concept study, we investigate an LLM's performance in screening relevant articles for a SLR with the novel screening of title and abstracts, reevaluation, and full-text review (STARR) protocol.
Methods
ENTGPT (based on GPT-4o) was compared to two human reviewers in article inclusion/exclusion decisions using the traditional and STARR screening protocols. ENTGPT was provided with inclusion and exclusion criteria, titles, abstracts, and full texts (if available) of the 850 articles retrieved in the original search. The model's decisions were compared to those made by two human reviewers.
Results
ENTGPT, using the STARR protocol, achieved 99.87% accuracy in article classification compared with human reviewers (95% CI: 0.99-1.0), including 100% specificity and 95% sensitivity. When using the traditional protocol, sensitivity declined to 35%. ENTGPT, using the traditional protocol, achieved 99.47% accuracy in article classification compared with human reviewers (95% CI: 0.99-1.0), including 100% specificity and 35% sensitivity. When using the STARR protocol, accuracy improved to 99.87% and sensitivity markedly increased to 95%.
Conclusions
ENTGPT accurately replicated human reviewers in article selection and data extraction for an otolaryngology SLR using the STARR and traditional protocols. This performance suggests that LLMs could be employed to significantly streamline the SLR process, potentially saving substantial time and resources for researchers.
LEVEL OF EVIDENCE
N/A.
Highlights
Large Language Models (LLMs) enable automation of initial abstract screening in systematic reviews, significantly reducing manual workload.
The effectiveness of LLMs heavily depends on prompt engineering, which must clearly translate inclusion and exclusion criteria into actionable instructio...
Shatskiy Alexander S., Ehab M. Deigheidy, S. E. Masyutina et al.· Complex Issues of Cardiovasc...· 0 citations
With targeted methodological refinements - including standardization of QUIPS implementation and validation against expert ratings - automated ROB assessments may meaningfully reduce time and cost of systematic reviews of prognosis studies in neurology and beyond.
S. Kampman, K. Braun, W. Otte· medRxiv· 0 citations
An enterprise LLM, prompted in POEM style, produced accurate, low-error clinical summaries that matched or exceeded expert-edited POEMs and were generally preferred by reviewers, though further research is needed to assess broader applicability and impact.
Richard Guthmann, Robert Martin, Erin Lee et al.· Journal of the American Boar...· 0 citations
LLM-based patient-facing pathology report interpretation shows potential to bridge specialist pathology language and patient communication, and evaluation should extend beyond readability to include fidelity to the original pathology report, patient understanding, safety, and usability.
Chen Wang, Jie Hao, Si-Jia Zhang et al.· International Journal of Med...· 0 citations
It is suggested that current LLMs can support literature screening, but their reliability depends strongly on the conceptual clarity of the review task and the structure of eligibility criteria, as well as practical implications for the transparent and responsible use of LLMs in informetric research and systematic revi...
Ni Cheng, Heng Dong, Xuan Han et al.· Aslib Journal of Information...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.