LLM-assisted writing and citation advantage: evidence from scientific publications before and after ChatGPT release
Abstract
Generative artificial intelligence has become a routine part of academic writing. While much of the debate has focused on questions of integrity and authorship, less attention has been paid to how AI-assisted writing may affect research evaluation itself. This paper asks a straightforward but important question: does the use of LLMs in academic writing change how scientific work is rewarded once it enters the citation system, even when the underlying research is legitimate? We study this question using a large-scale dataset that combines arXiv records with bibliometric information from Crossref, covering 234,073 research papers published between 2019 and 2025. To identify likely AI involvement, we train a text classification model on pre-LLM human-written abstracts and AI-generated rewrites, allowing us to detect stylistic features associated with LLM-assisted writing. We then examine how predicted LLM usage relates to citation outcomes using two complementary strategies: regression models with rich controls and fixed effects, and a counterfactual machine learning approach that compares observed citations to levels predicted from 2021–2022 citation patterns. Across both analyses, papers whose abstracts show signs of LLM-assisted writing receive more citations than otherwise similar papers. Specifically, papers classified as potentially written with the aid of LLMs exhibit approximately 7.04% higher citation counts on average (p < 0.01) relative to non-LLM papers. This difference remains after accounting for journal characteristics, collaboration size, reference lists, abstract length, and authors’ prior citation records, and it exceeds what would be expected based on historical citation dynamics alone. The results suggest that AI-assisted writing confers a modest but systematic advantage in visibility and discoverability that translates into higher measured impact. Rather than pointing to misconduct, these findings highlight a structural shift in how research is evaluated. As writing itself becomes easier to optimize through AI, citation-based indicators increasingly capture differences in presentation as well as differences in scientific contribution. The study therefore calls for greater caution in interpreting bibliometric measures in an academic environment where text is no longer produced solely by human authors.