Skip to content
Review Open access

An auditable evidence compiler for large language model-assisted systematic reviews

Sep 2026 · medRxiv · 0 citations
Medicine

TL;DR

This case provides a framework for inspecting synthesized evidence together with its provenance, statistical meaning and correction history and provides a framework for inspecting synthesized evidence together with its provenance, statistical meaning and correction history.

Abstract

Background: Large language models (LLMs) can support systematic reviews, but accurate individual outputs do not establish whether the final synthesis preserves the clinical question, accounts for statistical dependence and incorporates corrections. Objective: To develop and evaluate a framework linking LLM-assisted evidence processing to a versioned, auditable release of synthesis outputs. Methods: We used LLM agents to interpret sources and extract data. Deterministic code enforced statistical rules; investigators resolved material ambiguities and authorized release. We specified ten release properties covering evidence identities, statistical contributions and propagation of corrections. We retrospectively evaluated six integrity domains and historical failure events in one registered prognostic review, without an external comparator or held-out domain. Results: Fifty distinct root-cause events were documented, including 12 that had changed a pooled result before correction. Forty-six were resolved, and four remained disclosed limitations. The corpus comprised 454 reports, 445 studies, 441 cohort entities and 421 dependence clusters. Forty-one of 49 registered analyses were fitted, and eight retained explicit non-fitted states. All 39 source records across five principal analysis families reached a terminal source state. Two implementations within the project agreed across 1,217 numerical comparisons. All 94 file comparisons between release and publication packages were byte-identical. Two reviewers confirmed 39 principal records after seeing the same recommendations. Conclusions: This case provides a framework for inspecting synthesized evidence together with its provenance, statistical meaning and correction history. Comparative validity, generalizability and benefit in patient-centred care require independent evaluation.

Read PDF

Similar papers

PRISMA-LLM: An Empirical Reporting Framework for AI-Assisted Systematic Reviews

This work analyzes SciLitBench, a corpus of 888 review-automation papers with 14,726 annotations, to characterize changes in methods, review-stage use, evaluation and reported limitations, and introduces PRISMA-LLM, an empirically grounded framework separating implementation disclosure from consequence-sensitive evalua...

Miguel Zabaleta, Bai-Han Lin · 0 citations
Review Open access Oct 2026

Title and abstract screening for systematic reviews with Jev, a System One model: comparison with generative large language models

Large language models (LLMs) screen titles and abstracts without review-specific training, but generating screening decisions as text takes processing time and incurs API charges. We evaluated Jev, a non-generative model returning classification probabilities, on 4527 records from two systematic reviews of bipolar diso...

K. Matsui, Y. Takaesu · 0 citations
#artificial intelligence Preprint Sep 2026

CLEAR: Cross-Source Evidence Adjudication for Large Language Models in Medicine

CLEAR is proposed, an agentic framework for cross-source evidence adjudication in LLMs in medicine that independently generates candidate answers from three complementary pathways---parametric knowledge, locally curated corpora, and dynamically retrieved evidence---reflecting three common sources of information availab...

Shuai Wang, Yi-Ze Zhao, Qing-Yu Chen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.