Skip to content
Preprint

Beyond Static and Linear: What Attention Constraints Best Fit Human Reading Times?

Aug 2026 · 0 citations · 56 references
Computer Science

TL;DR

A systematic comparison of multiple attention-based memory mechanisms across different model sizes and training corpora finds that constraints that are sensitive to the content of intervening tokens consistently achieve the highest alignment with human reading times, outperforming distance-based constraints.

Abstract

Transformer-based language models are widely used as models of human language processing, yet their attention mechanisms allow lossless access to the full preceding context, unlike the limited memory systems of humans. We hypothesize that installing memory constraints into transformers'attention mechanisms can improve their fit to human behavioral data. While previous work has explored individual constraints in isolation, we conduct a systematic comparison of multiple attention-based memory mechanisms across different model sizes and training corpora, evaluating both psychometric predictive power for human reading times and grammatical competence. We additionally compare static constraints, in which the constraint strength is fixed throughout training, to dynamic memory curricula. We find that constraints that are sensitive to the content of intervening tokens consistently achieve the highest alignment with human reading times, outperforming distance-based constraints. We observe a dissociation between psychometric fit and grammatical competence under dynamic memory curricula, suggesting that Transformers cannot serve as a one-size-fits-all cognitive model.

View source

Similar papers

Open access 2026

Modeling Memory Effects in a Head-Final Language With Category Locality

Memory limitations have been assumed to be a major factor shaping human sentence comprehension, but characterizing memory-demanding structures in a broad-coverage and cross-linguistic manner has remained a challenge. A recent study suggested that such a general characterization can be obtained by assuming efficient c...

Shinnosuke Isono, Kohei Kajikawa, Yohei Oseki et al. · 2 citations
#artificial intelligence Preprint Sep 2026

For Your Eyes Only: Evaluating Coordination Between Isolated Language Model Instances

As model-generated content is increasingly consumed by other model instances in automated workflows, a practically important question arises: can a model embed a signal in natural language that an independent instance of the same model can detect, relying only on shared pre-training and task instructions, without any s...

Alexander Shirnin, A. Kudelya · 0 citations
#natural language process... Preprint Sep 2026

From Retrieval to Weights: Parametric Individualization of Small Language Models with Individual Text Corpora

We approach a cognitive simulation perspective on episodic and semantic memory in multiple-choice question answering by incorporating text from individual text corpora (ITC) into retrieval-augmented generation and DoRA fine-tuning. We web-crawl the search histories of 515 participants who answered 36 multiple-choice kn...

Christoph Wigbels, Ali Abusaleh, M. T. Jansen et al. · 0 citations
Jul 2026

Instruction-Tuned Models Locally Reuse Human Syntax More Than Humans Do

Syntactic convergence (the tendency of speakers to adapt in language towards the grammatical profiles of their interlocutors) is a well-documented feature of human dialogue widely considered to operate below conscious awareness. Whether large language models exhibit analogous syntactic convergence toward human users re...

Zandi Eberstadt · 0 citations
Preprint Aug 2026

Skill Issue: Are Skills Language-Invariant in LLMs?

This work quantifies cross-lingual skill inconsistency orthogonally from knowledge and general benchmark performance via multilingual self-play, and shows that skill discrepancies are a measurable major roadblock in the development of truly multilingual models.

Bobby Cheng, Adam Gaber, Zhengzhe Liu et al. · 1 citation
Jul 2026

RUMBA: Russian User Memory Benchmark

RUMBA (Russian User Memory BenchmArk) is introduced - a new benchmark for long-term conversational memory that provides a fine-grained taxonomy of memory-centric question types and a unified methodology accounting for semantic type, session scope, temporal reasoning, and the explicitness of temporal expressions.

E.D. Shevtsova, Inna Glebkina, Mark Baushenko et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.