Skip to content
Preprint

Disentangling the Expressivity of RoPE

Aug 2026 · 0 citations · 42 references
Computer Science

TL;DR

It is found that, if every rotary component is periodic, RoPE transformers recognize exactly the languages definable in past temporal logic with modular predicates.

Abstract

Two accounts recur in explanations of the success of rotary position embeddings (RoPE). Expressivity studies associate periodic position information with modular predicates, whereas mechanistic and long-context studies emphasize positional anchors and local offsets. We formalize both accounts for fully uniform, finite-precision soft-attention transformers. We find that, if every rotary component is periodic, RoPE transformers recognize exactly the languages definable in past temporal logic with modular predicates. Conventional RoPE is different: The rotations it computes never repeat. This yields a precision-dependent bounded simulation of fixed-offset look-back operators, rather than an all-length modular characterization. Controlled experiments match this separation: Constructed periodic schedules length-generalize on modular languages, while conventional RoPE behaves more like a bounded locality bias and can impair tasks requiring position-invariant access to distant context. Altogether, our findings shed light on RoPE transformers, bringing theoretical expressivity characterizations closer to models used in practice.

View source

Similar papers

Preprint Aug 2026

Protoreasoning in Tiny Transformers

We show that tiny transformers can profitably employ a simple form of Chain of Thought, which we call protoreasoning, allowing us to study step-by-step reasoning on ~1M-parameter models and opening up opportunities for much more detailed experimentation and analysis than is feasible for larger models. Current Large Lan...

E. Valle, Fergal Reid · 0 citations
Jul 2026

A Compositional Theory of Causally Masked Transformers

What types of decision problems can a causally masked, finite-precision transformer solve for inputs of arbitrary length? Existing answers often rely on idealized arithmetic, but under finite precision, rounding and evaluation order can change what information attention retains and therefore what the model can compute....

Franz Nowak, Ryan Cotterell, Reda Boumasmoud · 0 citations
#natural language process... Preprint Sep 2026

MeRoTune: RoPE-Safe Merging with a Tunable Dial

When you merge two fine-tuned models from the same base checkpoint by simply averaging their weights, you implicitly assume their attention subspaces are still aligned. Recent work attempts to fix misalignments by learning an invertible correction matrix, $M$, for each model's query and key projections. This correction...

Salman Faroz · 0 citations
#natural language process... Preprint Sep 2026

Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs

Reasoning in large language models unfolds through diverse functional operations, such as problem formulation, goal decomposition, and deduction. Although these operations are explicitly distinguished in text, little is known about how they are geometrically organized in representation spaces. To this end, we investiga...

Seogyeong Jeong, Jaehui Hwang, Dongyoon Han et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.