Skip to content
Review

Position Encoding in Transformers: From Absolute and Relative Methods to Rotary Position Embeddings and Long-Context Scaling

Aug 2026 · 0 citations · 25 references
Computer Science

TL;DR

A central conclusion is that the ability to compute positional features beyond the training length does not imply reliable long-context generalization; context extension must be evaluated through short-context retention, position-wise perplexity, retrieval, reasoning, and long-context code tasks.

Abstract

Self-attention models content-dependent interactions between tokens but does not by itself encode token order. Position encoding addresses this limitation by introducing absolute coordinates, relative distances, or position-dependent rotations into Transformer representations and attention scores. This technical survey develops a unified account of sinusoidal and learned absolute position embeddings, Shaw-style relative position representations, Transformer-XL, T5 relative position bias, ALiBi, and Rotary Position Embeddings (RoPE). We derive how RoPE converts absolute position indices into relative phase differences in Query-Key inner products and compare these methods in terms of where position is injected, computational cost, compatibility with KV caching, and length extrapolation. We then examine long-context extensions, including Position Interpolation, RoPE scaling laws, NTK-aware scaling, Dynamic NTK, NTK-by-parts, YaRN, LongRoPE, and LongRoPE2, with emphasis on frequency allocation, attention rescaling, training length, and target context length. We also summarize implementation considerations, evaluation protocols, and position-encoding choices in representative large language models. A central conclusion is that the ability to compute positional features beyond the training length does not imply reliable long-context generalization; context extension must be evaluated through short-context retention, position-wise perplexity, retrieval, reasoning, and long-context code tasks.

View source

Similar papers

#machine learning Preprint Sep 2026

Content-Based Addressing for Long Context

It is proved that this construction preserves local RoPE exactly, leaves the attention comparison between two fixed tokens unchanged when other units are inserted or reordered, and does not create new relative rotations merely because more units are added.

M. Godavarti · 0 citations
#natural language process... Preprint Sep 2026

Distance generalization in transformers: why bother with positional encoding?

A thorough investigation of distance generalization probes performance when inter-token distances are changed between training and inference, while keeping a fixed context length, finding that it is paramount to improve the understanding of the underlying mechanisms.

D. Nevermann, Claudius Gros · 0 citations
#artificial intelligence Preprint Aug 2026

Higher-Dimensional Rotary Position Embedding

HDR-RoPE is proposed, which extends RoPE from independent 2D rotations to higher-dimensional rotations and introduces a Paley-I orthogonal basis to obtain balanced, isotropic, and dense phase mixing within each rotation subspace and significantly enhances channel coupling and rotational degrees of freedom while maintai...

Yixing Li, Ruo-Bing Xie, Yu-Dong Zhang et al. · 0 citations
Preprint Aug 2026

ZetaGPT: A Reference Implementation of Positional--Encoding--Free State--Space--Attention Language Models

This work explores a language model architecture that integrates causal state-space equations to implicitly encode positional information before attention computation, and establishes a compact, reproducible reference implementation for the development and empirical study of positional-encoding-free language models.

Róisín Luo · 0 citations
#natural language process... Preprint Sep 2026

Paragraph Boundaries Are Not White Space:Compression Depth as the Signature of Hierarchical Structure

A hierarchical rotary positional encoding (hRoPE) that represents paragraph, sentence, and token indices as separate channels, hold the token sequence fixed, intervene on the paragraph coordinate $p_1$, and measure cross-paragraph attention with a token-distance-exact estimator is used.

Shu-Yang Xiang · 0 citations
Book Open access Aug 2026

Dynamic Positional Attention Modulation for Parameter-Efficient Fine-Tuning of Large Language Models

DyPAM (Dynamic Positional Attention Modulation) is proposed, a PEFT method that adapts how positional information contributes to attention by operating directly on the query and key representations, performing fine-grained adaptation of positional attention aligned with the RoPE-induced structure without modifying the...

Dayan Pan, Jing-Yuan Wang, Xie Yu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.