It is found that LLMs systematically reinforce the Matthew effect by favoring highly cited papers, with field-specific variation in the rate at which generated references match real papers in bibliometric databases.
Abstract
The spread of scientific knowledge depends on how researchers discover and cite prior work. Large language models (LLMs) now add a new layer to this process, but their alignment with human citation practices across domains remains unclear. Here, we compare human citations with GPT-4ogenerated reference suggestions produced from paper metadata and abstracts. Analyzing 274, 951 generated references for 10, 000 focal papers, we find that LLMs systematically reinforce the Matthew effect by favoring highly cited papers, with field-specific variation in the rate at which generated references match real papers in bibliometric databases. Generated references diverge from groundtruth reference lists by favoring more recent papers, shorter titles, and smaller author teams. Yet they remain semantically aligned with focal-paper content at levels comparable to human references, reproduce similar local citation-network structure, and reduce author self-citations. These results show that LLMs can generate content-relevant bibliographic suggestions from parametric knowledge alone, but that they also amplify dominant citation patterns. As such tools become routine in research workflows, they may reshape how scientific communities discover, prioritize, and build on prior work.
This study presents one of the first comprehensive evaluations of multiple state-of-the-art (SOTA) large language models (LLMs) for citation function classification, achieving new SOTA results on the ACL-ARC dataset.
Daniel Vodička, Jakub Šmíd, Pavel Král et al.· International Conference on...· 0 citations
The Knowledge Contribution Taxonomy (KCT), derived from the Scientific Research Logic Model, is proposed, which identifies the type of knowledge a cited paper contributes based on the citation context, and classifies citations into Method, Resource Tool, Empirical Finding, and Background, further distinguishing core fr...
Zhibang Quan, Zhentao Liang, Ming Ma et al.· 0 citations
Scientific citations carry rhetorical intent. Scholars may cite prior work positively (supporting), negatively (contrasting), or neutrally (mentioning). As large language models (LLMs) increasingly assist scientific writing, whether they reproduce citations with the same rhetorical intent as humans remains unclear. We...
Yixuan Liu, Lin Chen, Zhuoqi Liu et al.· 0 citations
Impact-DPO integrates temporally informed prompting with direct preference optimization, enabling LLMs to learn comparative influence patterns without explicit graph message passing, and formalizes citation forecasting as pairwise preference learning on temporal text-attributed graphs.
Parham Hamouni, Ebrahim Bagheri· ACM Transactions on Intellig...· 0 citations
Scientific discovery depends on finding prior literature that shapes what comes next. Existing retrieval systems optimize for relevance and popularity, often favoring central papers over less familiar works that later prove generative. We introduce \textbf{MUSES}, a million-instance benchmark for prospective intellectu...
Rohan Pandey, Sunjae Kwon, Hong Yu· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.