The Predictive Validity of Peer Review for Article Impact: A Systematic Review and Meta-Analysis
Abstract
Objective: Peer review serves as the primary quality assurance mechanism in scholarly publishing, yet its capacity to forecast the downstream scientific impact of published research remains inadequately characterized. This study systematically reviews and meta-analyzes empirical evidence examining the association between peer review evaluations and post-publication article impact, measured through citation counts and altmetric attention. Methods: A systematic literature search was conducted across PubMed, Web of Science, Scopus, PsycINFO, and Embase from database inception to October 2025, following PRISMA 2020 guidelines. Studies reporting quantitative associations between peer review scores and post-publication impact metrics were included. A random-effects meta-analysis was performed to pool correlation coefficients, with heterogeneity assessed using the I² statistic and publication bias evaluated via funnel plots and Egger’s regression test. Quality assessment employed a modified Newcastle-Ottawa Scale. Subgroup analyses examined discipline, journal impact factor, and review model. Bayesian meta-analysis with informative priors was conducted to complement frequentist estimates and formally evaluate evidence for the null hypothesis. Results: Of 1,247 identified records, 12 studies met inclusion criteria, encompassing 8,571 manuscripts across biomedical and social science disciplines. The pooled correlation between peer review scores and two-year citation counts was weak and non-significant ( r = 0.04, 95% CI: -0.02 to 0.10, p = 0.21; I² = 89%). Bayesian analysis yielded a posterior estimate of r = 0.04 (95% CrI: -0.03 to 0.11) with a Bayes Factor of 0.23, indicating moderate evidence favoring the null hypothesis relative to a small positive effect. Subgroup analyses and meta-regression revealed no significant moderation by journal impact factor, academic discipline, review model, number of reviewers, or reviewer experience. A moderate positive association emerged between review scores and Altmetric Attention Scores ( r = 0.28, 95% CI: 0.15 to 0.40, p < 0.001; I² = 72%). Conclusions: Traditional peer review scores demonstrate negligible predictive validity for citation-based scholarly impact. The modest correlation with altmetrics suggests peer review may capture dimensions related to public engagement rather than conventional scholarly influence. These findings challenge foundational assumptions about peer review as a predictive quality filter and support calls for reformed research evaluation frameworks, consistent with the San Francisco Declaration on Research Assessment.