Skip to content

Distance generalization in transformers: why bother with positional encoding?

Sep 2026 · 0 citations · 26 references
Computer Science

TL;DR

A thorough investigation of distance generalization probes performance when inter-token distances are changed between training and inference, while keeping a fixed context length, finding that it is paramount to improve the understanding of the underlying mechanisms.

Abstract

Out-of-distribution length generalization, namely to extrapolate a task from short to longer context, has been studied intensively for transformers. Here we focus on distance generalization, which probes performance when inter-token distances are changed between training and inference, while keeping a fixed context length. We construct two synthetic delay copy tasks, both involving finite distances between source and recall, where tokens are copied either fully or selectively, and test models on delays unseen during training. We address three questions: (A) Do positional encoding schemes such as RoPE and ALiBi improve distance resolution relative to no positional encoding (NoPE)? (B) How does data diversity, the number of inter-token distances seen in training, affect performance? (C) When is distance transfer learning positive or negative? We present a thorough investigation, finding that it is paramount to improve our understanding of the underlying mechanisms.

View source

Similar papers

Review Aug 2026

Position Encoding in Transformers: From Absolute and Relative Methods to Rotary Position Embeddings and Long-Context Scaling

A central conclusion is that the ability to compute positional features beyond the training length does not imply reliable long-context generalization; context extension must be evaluated through short-context retention, position-wise perplexity, retrieval, reasoning, and long-context code tasks.

Jiguo Li · 0 citations
Preprint Aug 2026

TANGO: Treating Tokens as Operators

Transformers separate cross-token mixing in self-attention from token-wise transformation in feed-forward networks. We ask whether combining these operations can lower predictive loss under fixed data and parameter budgets. To do so, we introduce the Token-Aggregated Nonlinear Gating Operator (TANGO) model. TANGO compu...

Joshua Nunley · 0 citations
#artificial intelligence Preprint Sep 2026

How Local Mixing Encodes Relative Position in Global NoPE Attention

An explanation of how hybrid models of this sort can implicitly encode position at global NoPE layers is developed, and insights are provided for how to encode position in a way that can extrapolate to longer sequence lengths indefinitely.

Cutter Dawes, Nick Alonso, Tomas Figliolia et al. · 0 citations
#machine learning Preprint Sep 2026

The Token Before the Value Is the Key: How Hybrid Architectures Organize Induction Circuits

Hybrid language models can improve capability as well as efficiency, raising the question of how architectural complementarity becomes learned computation. We examine the established induction roles of Carrying predecessor information, Matching a source by content, and Copying its value. How are these position-sensitiv...

Ke Cheng, Xin Xu, Yi-Xiao Chen et al. · 0 citations
#machine learning Preprint Sep 2026

RoPE is Dead, Long Live RoPE: Towards Scalable Data-aware Positional Encodings

Transformers process tokens without any inherent notion of order, making positional encoding a fundamental requirement rather than an architectural refinement. Rotary Position Embedding (RoPE) has become the default positional encoding in modern language models, yet it is heavily biased toward nearby tokens. Existing a...

Jarod Lévy, Mathurin Videau, Jad Yehya et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SMat-Attention: Structured Long-Context Sequence Modeling

Long-context sequence models face a fundamental tradeoff: softmax attention uses flexible token-level interactions at quadratic cost, whereas linear attention obtains linear-time training and constant-time decoding by compressing history into a fixed-size state. In this work, we ask whether we can connect these regimes...

E. Anand, Abdullah Ateyeh, Archer Wang et al. · 2 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.