Skip to content
Open access

HOW DEEP DO LARGE LANGUAGE MODELS INTERNALIZE SCIENTIFIC LITERATURE AND CITATION PRACTICES?

Jul 2026 · Quantitative Science Studies · 2 citations

TL;DR

It is found that LLMs systematically reinforce the Matthew effect by favoring highly cited papers, with field-specific variation in the rate at which generated references match real papers in bibliometric databases.

Abstract

The spread of scientific knowledge depends on how researchers discover and cite prior work. Large language models (LLMs) now add a new layer to this process, but their alignment with human citation practices across domains remains unclear. Here, we compare human citations with GPT-4ogenerated reference suggestions produced from paper metadata and abstracts. Analyzing 274, 951 generated references for 10, 000 focal papers, we find that LLMs systematically reinforce the Matthew effect by favoring highly cited papers, with field-specific variation in the rate at which generated references match real papers in bibliometric databases. Generated references diverge from groundtruth reference lists by favoring more recent papers, shorter titles, and smaller author teams. Yet they remain semantically aligned with focal-paper content at levels comparable to human references, reproduce similar local citation-network structure, and reduce author self-citations. These results show that LLMs can generate content-relevant bibliographic suggestions from parametric knowledge alone, but that they also amplify dominant citation patterns. As such tools become routine in research workflows, they may reshape how scientific communities discover, prioritize, and build on prior work.

Read PDF

Similar papers

Review Jul 2026

Large Language Models for Citation Function Classification

This study presents one of the first comprehensive evaluations of multiple state-of-the-art (SOTA) large language models (LLMs) for citation function classification, achieving new SOTA results on the ACL-ARC dataset.

Daniel Vodička, Jakub Šmíd, Pavel Král et al. · 0 citations
Preprint Aug 2026

Data Citation for Large Language Models: A Challenge

It is argued that data citation for large language models is an open challenge, distinct from document-level citation grounding and harder to solve.

G. Silvello · 0 citations
Preprint Aug 2026

From citation intent to knowledge contribution: Classifying what cited papers actually contribute

The Knowledge Contribution Taxonomy (KCT), derived from the Scientific Research Logic Model, is proposed, which identifies the type of knowledge a cited paper contributes based on the citation context, and classifies citations into Method, Resource Tool, Empirical Finding, and Background, further distinguishing core fr...

Zhibang Quan, Zhentao Liang, Ming Ma et al. · 0 citations
#natural language process... Preprint Sep 2026

Citing Less Critically: LLMs Reshape the Rhetoric and Reach of Scientific Citation

Scientific citations carry rhetorical intent. Scholars may cite prior work positively (supporting), negatively (contrasting), or neutrally (mentioning). As large language models (LLMs) increasingly assist scientific writing, whether they reproduce citations with the same rhetorical intent as humans remains unclear. We...

Yixuan Liu, Lin Chen, Zhuoqi Liu et al. · 0 citations
Open access Aug 2026

Predicting Scholarly Impact with Temporal Preference Alignment

Impact-DPO integrates temporally informed prompting with direct preference optimization, enabling LLMs to learn comparative influence patterns without explicit graph message passing, and formalizes citation forecasting as pairwise preference learning on temporal text-attributed graphs.

Parham Hamouni, Ebrahim Bagheri · 0 citations
Preprint Aug 2026

MUSES: A Benchmark for Prospective Intellectual-Roots Retrieval

Scientific discovery depends on finding prior literature that shapes what comes next. Existing retrieval systems optimize for relevance and popularity, often favoring central papers over less familiar works that later prove generative. We introduce \textbf{MUSES}, a million-instance benchmark for prospective intellectu...

Rohan Pandey, Sunjae Kwon, Hong Yu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.