Skip to content

Author

Benjamin K. Bergen

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Review Open access 2026

Large Language Models as Distributional Baselines for Language Tasks

In order to ask questions about the mechanisms underpinning human cognition, researchers must control for properties of stimuli that could confound detected effects. In experiments involving linguistic stimuli, this includes properties like frequency, length, and neighborhood size of those stimuli, which are known to affect behavioral and neural responses. With improvements in the performance and usability of language models, it is now possible to also control for how predictable stimuli and their parts are, on the basis of the distributions of words alone: their distributional predictability. This coincides with a resurgence of interest in the possibility that statistical language learning may underlie a broad range of human cognitive phenomena; indeed, there are both theoretical and empirical reasons to believe that humans rely on distributional information during certain cognitive tasks. This creates a confound, whereby experimental operationalizations of psychological constructs with linguistic stimuli may not in fact be testing what they are intended to test. Thus, the central contributions of this paper are twofold: first, we articulate the conditions under which distributional predictability threatens the internal validity of an experiment; and second, we provide concrete recommendations for how to control for this potential confound. Beyond these primary contributions, we survey techniques for measuring distributional predictability, review theoretical and empirical work supporting the role of distributional statistics in human cognition, and present several case studies illustrating the range of possible outcomes—from the “distributional baselines” only marginally affecting theoretical inferences to constituting fully deflationary confounds. We also enumerate and address potential objections to this approach. This paper is primarily intended for researchers in psychology, cognitive science, and linguistics who use linguistic stimuli but have not yet incorporated distributional baselines into their work.

Sean Trott, James A. Michaelov, Cameron R. Jones et al. · 0 citations