Skip to content

Scaling Interpretable Transformers with Parity Bottleneck Layers

Jul 2026 · arXiv.org · Vol abs/2607.20652 · 0 citations · 50 references
Computer Science

TL;DR

The ParityTransformer is introduced, a GPT-2-scale architecture whose intermediate representations are efficient and wide / sparse by design, and seen as a step toward training models whose internal representations are interpretable by design rather than recovered post hoc.

Abstract

Language models are thought to exhibit the phenomenon of superposition, representing many more features than dimensions in their residual streams. Sparse autoencoders (SAEs) are designed to recover such features post-hoc, but training models that are interpretable by construction has remained impractical, as a per-layer over-complete bottleneck is prohibitively expensive in both memory and compute. To overcome this issue, we introduce the ParityTransformer, a GPT-2-scale architecture whose intermediate representations are efficient and wide / sparse by design. At each layer, a Deep Parity Bottleneck (DPB) replaces a learned over-complete basis with a parameter-free algebraic dictionary, providing a deterministic incoherence guarantee and eliminating the memory requirements that have prevented per-layer interpretable bottlenecks at scale. A DPB is a hierarchically structured sparse bottleneck which efficiently enforces sparsity using a multi-level mixture-of-experts approach: a hardware-aware implementation that closes the cost gap between activation sparse and dense training to a manageable interpretability tax. Empirically, ParityTransformers perform at least as well as post-hoc SAEs on sparse probing tasks, while out-performing on measures of feature absorption, steering effectiveness, and fine-grained causal interventions. Because subsequent computation acts only on features that survive the sparse bottleneck, the ParityTransformer's features are native to the model's forwards pass by construction, addressing the question of whether SAEs probe features the model actually uses during computation. We see this as a step toward training models whose internal representations are interpretable by design rather than recovered post hoc.

View source

Similar papers

Preprint Aug 2026

Finding Usable Weight Mechanisms with Tiled SVD

This work proposes extracting mechanism mounts directly from linear sites by column-tiled SVD: each mount is a triple read as trigger, write, and strength, and evaluates mounts with a pre-registered suite judged on full-write energy lift rather than tile-local lift.

Ash Manvi, Samreena Tajreen · 0 citations
Preprint Aug 2026

The Sparsity Whisperer

A family of difference-informed pruning methods built upon this principle are introduced, suggesting that preserving output differences is a broadly useful and composable signal for post-training LLM sparsification.

Linghao Kong, Inimai Subramanian, Micah Adler et al. · 0 citations
#machine learning Preprint Sep 2026

Manifold Projection and Iterative Autoencoder Refinement for Masked Language Modeling

This work replaces attention with a stack of autoencoder-based mixing modules, one operating over local neighborhoods, one over the full sequence, and one across attention heads, each compressing and reconstructing its input through a bottleneck, and its width is a hyperparameter rather than a training effect.

Narges Mokhtari, F. Haddadi, Ebrahim Rezaii · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.