Skip to content
Preprint

A scaling law of contextual persistence in human language

Jul 2026 · 0 citations · 42 references
Computer Science

Abstract

Human language exhibits lawful structure at the level of words (frequency, vocabulary growth) and word pairs (co-occurrence across distance). Here we show that the arrangement of words in sequence -- a central determinant of meaning -- obeys a comparable law. Using large language models as probabilistic probes, we measured the reduction in target perplexity conferred by prior context at distance d beyond that of the same words scrambled; this difference, the contextual persistence function P(d), isolates the influence of arrangement. Across ten corpora spanning six language families and written and spoken modalities, P(d) decayed approximately as 1/d ($P(d) \propto d^{-\alpha}$, mean $\alpha = 1.04$; median $r^2 = 0.96$). The effect vanished in scrambled and synthetic controls, replicated across independent probes, and did not appear in genomic or protein sequences under domain-native models. An exponent near 1 distributes contextual influence approximately uniformly across logarithmic timescales. The results establish a scaling law of contextual persistence in human language.

View source

Similar papers

Preprint Aug 2026

From Exposure to Expectation: Frequency, Surprisal, and Language Across Development in Spanish

Surprisal, the negative log-probability a language model assigns to a word given its preceding context, reliably predicts adult reading times. Does it contribute as much to explaining when children acquire individual words? Frequency reflects a learner's cumulative exposure to a word, whereas surprisal reflects how predictable a single occurrence is given its context. We investigate this question across two corpus-based studies of Spanish. In Study 1, we modeled age of acquisition (AoA) for 225 Spanish nouns using lexical frequency and contextual diversity from child-directed speech, plus surprisal from three language models differing in architecture and training language (BETO, BERTIN, mGPT). Frequency strongly predicted AoA (r=-.597, p<.001); surprisal added little beyond frequency and word length, including in a naturalistic-context analysis. In Study 2, we modeled adult fixation durations in the Chilean Spanish subsample of the Multilingual Eye-movement Corpus (MECO Wave 2), using mGPT surprisal alongside two independent frequency measures. Surprisal robustly predicted longer fixation durations after controlling for frequency and word length, consistent across both frequency sources. A matched word-type-level comparison showed the surprisal-behavior association was stronger in reading than in acquisition (z=3.63, p<.001). The findings suggest cumulative lexical exposure and contextual predictability play different roles across the language trajectory: frequency is particularly informative about when early lexical representations are acquired, whereas surprisal captures moment-to-moment processing difficulty in an already-established linguistic system. We discuss this pattern in relation to usage-based and entrenchment-based accounts of lexical development and to the evaluation of language models as models of human language behavior.

Francisco López · 0 citations
Preprint Aug 2026

Local and Global Regimes of Geometric Complexity in Language Model Representations

A scale-dependent transition between two ID regimes is found: at low lexical diversity, conditions with fewer unique final words produce higher ID, while at high lexical diversity, this ordering reverses, and conditions with more unique words produce higher ID.

Arwa Osman, Marco Baroni, Iuri Macocco · 0 citations
Review Open access 2026

Large Language Models as Distributional Baselines for Language Tasks

In order to ask questions about the mechanisms underpinning human cognition, researchers must control for properties of stimuli that could confound detected effects. In experiments involving linguistic stimuli, this includes properties like frequency, length, and neighborhood size of those stimuli, which are known to affect behavioral and neural responses. With improvements in the performance and usability of language models, it is now possible to also control for how predictable stimuli and their parts are, on the basis of the distributions of words alone: their distributional predictability. This coincides with a resurgence of interest in the possibility that statistical language learning may underlie a broad range of human cognitive phenomena; indeed, there are both theoretical and empirical reasons to believe that humans rely on distributional information during certain cognitive tasks. This creates a confound, whereby experimental operationalizations of psychological constructs with linguistic stimuli may not in fact be testing what they are intended to test. Thus, the central contributions of this paper are twofold: first, we articulate the conditions under which distributional predictability threatens the internal validity of an experiment; and second, we provide concrete recommendations for how to control for this potential confound. Beyond these primary contributions, we survey techniques for measuring distributional predictability, review theoretical and empirical work supporting the role of distributional statistics in human cognition, and present several case studies illustrating the range of possible outcomes—from the “distributional baselines” only marginally affecting theoretical inferences to constituting fully deflationary confounds. We also enumerate and address potential objections to this approach. This paper is primarily intended for researchers in psychology, cognitive science, and linguistics who use linguistic stimuli but have not yet incorporated distributional baselines into their work.

Sean Trott, James A. Michaelov, Cameron R. Jones et al. · 0 citations
Review Open access Aug 2026

Unifying the structures of language in a neural population code.

It is concluded that explaining how language can emerge from neural population codes, in both biological and artificial systems, will not be achieved through the incremental refinement of algebraic-symbolic theories but will demand new theoretical paradigms.

Samuel A. Nastase, Zaid Zada, A. Goldberg et al. · 0 citations
Jul 2026

Language is a factor in the identification and processing of number words.

Since the 1990s, the fields of numerical cognition and psycholinguistics have evolved largely independently. However, the research since then demonstrates that language and numbers are intrinsically related, particularly in multilingual contexts. Language characteristics (e.g., morpho-syntactic properties) and language status (e.g., monolingual vs. multilingual) influence numerical processing at multiple levels, from number identification and retrieval to advanced mathematical skills. This article builds on the contribution of Frenck-Mestre and Vaid (1992), "Language as a Factor in the Identification of Ordinary Words and Number Words", to further elaborate how language contributes to the learning of numerical concepts, especially the acquisition and representation of number words across development. While their findings provided initial important insights into bilingual number processing, subsequent research, which we discuss in the current article, further shows that language and multilingual experiences can shape lexical, semantic, and lexico-semantic access to numbers. We finish by outlining some future research directions for the study of multilingual number representations.

Rémy Lachelin, Christine Schiltz, M. Marinova · 0 citations
Open access Aug 2026

Cognitive Processes of Probabilistic Prediction in Reading: Language Model Surprisal Across Model Sizes, Token Granularity, and Reading Paradigms

Surprisal, the negative log probability of a word given its context, is the dominant computational metric for quantifying reading difficulty and a common item difficulty estimator in reading research. Yet how the language model family, surprisal granularity, and corpus type jointly shape the surprisal–reading time link remains unclear. We conducted a secondary analysis of two public English reading corpora: the Natural Stories Corpus (self-paced reading, 181 readers) and the Provo Corpus (eye tracking with cloze norms). Surprisal was computed from GPT 2 (Small, XL), Llama 2 (7B, 13B), and Llama 3 (8B), together with a 5-gram baseline and human cloze norms. Linear mixed-effects models tested the baseline contribution of neural surprisal, inverse scaling within model families, word-level versus sub-word aggregation, and the linking function shape. Neural surprisal contributed reliable variance above a strong baseline in both corpora. A clear inverse scaling pattern emerged: GPT 2 Small produced the largest fit improvements, exceeding GPT 2 XL, Llama 2 13B, and Llama 3 8B. Word-level aggregation outperformed sub-word aggregation, especially for measures of early lexical access. Non-parametric analyses supported an approximately linear linking function, and cloze norms carried information not fully captured by neural surprisal. These findings show that larger models are not automatically better cognitive models of reading and that surprisal granularity is not a neutral analytic choice.

Shuting Liu, Yong Mei · 0 citations