Skip to content

Content-Based Addressing for Long Context

Sep 2026 · 0 citations · 6 references
Computer Science

TL;DR

It is proved that this construction preserves local RoPE exactly, leaves the attention comparison between two fixed tokens unchanged when other units are inserted or reordered, and does not create new relative rotations merely because more units are added.

Abstract

Rotary position embedding (RoPE) uses each token's integer position to determine the rotation applied inside attention. This works well for local token order, but increasing context length creates a positional train-test mismatch: RoPE produces relative rotations at offsets not seen during training. Methods that rescale, interpolate, randomize, or bias positions specify how attention handles those offsets, but still derive positional information from a growing token counter. We instead divide a token stream into units, retain ordinary RoPE positions within each unit, and assign every completed unit an address computed from its content. Adding units then applies the same learned map to new content rather than extending a positional range or an identifier table. We prove that this construction preserves local RoPE exactly, leaves the attention comparison between two fixed tokens unchanged when other units are inserted or reordered, and does not create new relative rotations merely because more units are added. In a character-level Tiny Shakespeare diagnostic, all-token validation perplexity remains approximately constant from contexts of 256 to 4096 characters. A second diagnostic shows that content-based addressing can retrieve and use information from multiple serialized facts. These are controlled shallow experiments, not scale benchmarks, but they support a direct prescription: use position to address locally and content to address across units.

View source

Similar papers

#natural language process... Preprint Oct 2026

WavePrune: One period is often enough for RoPE

Rotary Position Embedding (RoPE) encodes token positions by rotating each two-dimensional channel of the query and key vectors at a channel-specific frequency, making the attention logits invariant to a common shift of positions. However, this rotation is periodic, and it leads to position aliasing where relative posit...

Guan-Cheng Du, Luo-Tian Huang, Shao-Wen Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Aperture: Merge-Consistent Rotary States for Compressed Tokens

Aperture stores Fourier moments of the token's weighted support at the model's rotary frequencies, and it is proved that these moments have minimal real dimension among continuous states sufficient for the selected expected rotary interactions.

Yu-Hao Du, Shu-Nian Chen · 0 citations
#natural language process... Preprint Sep 2026

Paragraph Boundaries Are Not White Space:Compression Depth as the Signature of Hierarchical Structure

A hierarchical rotary positional encoding (hRoPE) that represents paragraph, sentence, and token indices as separate channels, hold the token sequence fixed, intervene on the paragraph coordinate $p_1$, and measure cross-paragraph attention with a token-distance-exact estimator is used.

Shu-Yang Xiang · 0 citations
#natural language process... Preprint Sep 2026

MeRoTune: RoPE-Safe Merging with a Tunable Dial

When you merge two fine-tuned models from the same base checkpoint by simply averaging their weights, you implicitly assume their attention subspaces are still aligned. Recent work attempts to fix misalignments by learning an invertible correction matrix, $M$, for each model's query and key projections. This correction...

Salman Faroz · 0 citations
#artificial intelligence Preprint Oct 2026

Later Is Better: Token Reduction for ViTs Under Distribution Shift

Training-free token reduction accelerates vision transformers by removing redundant tokens across layers, recovering most of the original accuracy at a fraction of the compute. These methods, however, are designed and evaluated primarily on clean data, and under real-world distribution shift their accuracy gap to the u...

Hyeong-Tae Cha, Hyungjun Yoon, Sung-Ju Lee · 0 citations
#machine learning Preprint Sep 2026

Analytic-Walk Rotary Positional Encodings for Graphs

Rotary position encodings make attention sensitive to relative position, but extending them to graphs requires choosing how graph structure enters the rotation. Previous works assign each node a rotation from spectral coordinates, so the rotary factor between two nodes depends only on their endpoints and cannot disting...

Jia-Qing Xie, Yu-Xin Wang, Xi-Peng Qiu · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.