Skip to content
Preprint

ZetaGPT: A Reference Implementation of Positional--Encoding--Free State--Space--Attention Language Models

Aug 2026 · 0 citations · 28 references
Computer Science

TL;DR

This work explores a language model architecture that integrates causal state-space equations to implicitly encode positional information before attention computation, and establishes a compact, reproducible reference implementation for the development and empirical study of positional-encoding-free language models.

Abstract

Transformer-based language models rely on self-attention, whose computation is permutation-equivariant and therefore lacks an intrinsic mechanism for representing token order. Existing architectures address this limitation by explicitly incorporating positional information through learned positional embeddings or hand-crafted positional encodings, such as rotary positional encoding (RoPE), treating positional information as an architecturally acquired capability rather than an inherent property of the model. Motivated by the pursuit of positional-encoding-free architectures, this work explores a language model architecture that integrates causal state-space equations to implicitly encode positional information before attention computation. Specifically, each model block applies a causal state-space equation before self-attention, allowing recurrent state dynamics to encode sequential information into token representations. Consequently, subsequent attention layers operate on position-aware representations without requiring explicit positional encodings while retaining the expressive modeling capacity of self-attention. We present \textsc{ZetaGPT}, a compact hybrid language model designed for research, rapid prototyping, algorithm verification, and educational applications. In addition to the proposed architecture, \textsc{ZetaGPT} provides a fully open-source, end-to-end training pipeline encompassing dataset construction, tokenizer training, pretraining, supervised fine-tuning, reinforcement learning from human feedback (RLHF), and chain-of-thought (CoT) reasoning via pure reinforcement learning. To the best of our knowledge, \textsc{ZetaGPT} is the first open-source small language model without explicit positional encoding and establishes a compact, reproducible reference implementation for the development and empirical study of positional-encoding-free language models.

View source

Similar papers

The Evolution of Attention in Large Language Models: Mechanisms, Trade-offs, and Emerging Trends

Self-attention gives LLMs fine-grained, query-dependent access to context, but dense token interactions incur quadratic prefill cost and a key--value cache growing with context length. Research thus spans explicit-memory compression, sparse access, recurrent state construction, structured state dynamics, and heterogene...

Zhen-Tao Tan, Jing-Yi Shen, Yan-Bo Li et al. · 0 citations
Review Aug 2026

Position Encoding in Transformers: From Absolute and Relative Methods to Rotary Position Embeddings and Long-Context Scaling

A central conclusion is that the ability to compute positional features beyond the training length does not imply reliable long-context generalization; context extension must be evaluated through short-context retention, position-wise perplexity, retrieval, reasoning, and long-context code tasks.

Jiguo Li · 0 citations
#artificial intelligence Preprint Sep 2026

How Local Mixing Encodes Relative Position in Global NoPE Attention

An explanation of how hybrid models of this sort can implicitly encode position at global NoPE layers is developed, and insights are provided for how to encode position in a way that can extrapolate to longer sequence lengths indefinitely.

Cutter Dawes, Nick Alonso, Tomas Figliolia et al. · 0 citations
#machine learning Preprint Sep 2026

The Token Before the Value Is the Key: How Hybrid Architectures Organize Induction Circuits

Hybrid language models can improve capability as well as efficiency, raising the question of how architectural complementarity becomes learned computation. We examine the established induction roles of Carrying predecessor information, Matching a source by content, and Copying its value. How are these position-sensitiv...

Ke Cheng, Xin Xu, Yi-Xiao Chen et al. · 0 citations
#natural language process... Preprint Aug 2026

All You Need Is Non-Commutative Words

It is shown that the noncommutativity of matrix product captures word order without positional encodings (PEs) and yields several capabilities, including antisymmetric self-attention with no query, key, or value projections, and parallel composition of variable-length text chunks at a reduced attention cost.

Carla Quispe Flores, Stanley Salvatierra, Renan Cabrera · 0 citations
Sep 2026

Information-Theoretic Analysis of Positional Encoding Strategies in Vision Transformers: A Comparative Study of Four Approaches.

Vision Transformers (ViTs) rely on positional encoding (PE) because self-attention has no native notion of token order or image-grid location, yet the information-theoretic properties of different PE strategies and their downstream consequences for model behaviour remain insufficiently characterised. We present a syste...

D. Bandur, M. Bandur, B. Jakšić · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.