The ParityTransformer is introduced, a GPT-2-scale architecture whose intermediate representations are efficient and wide / sparse by design, and seen as a step toward training models whose internal representations are interpretable by design rather than recovered post hoc.
Abstract
Language models are thought to exhibit the phenomenon of superposition, representing many more features than dimensions in their residual streams. Sparse autoencoders (SAEs) are designed to recover such features post-hoc, but training models that are interpretable by construction has remained impractical, as a per-layer over-complete bottleneck is prohibitively expensive in both memory and compute. To overcome this issue, we introduce the ParityTransformer, a GPT-2-scale architecture whose intermediate representations are efficient and wide / sparse by design. At each layer, a Deep Parity Bottleneck (DPB) replaces a learned over-complete basis with a parameter-free algebraic dictionary, providing a deterministic incoherence guarantee and eliminating the memory requirements that have prevented per-layer interpretable bottlenecks at scale. A DPB is a hierarchically structured sparse bottleneck which efficiently enforces sparsity using a multi-level mixture-of-experts approach: a hardware-aware implementation that closes the cost gap between activation sparse and dense training to a manageable interpretability tax. Empirically, ParityTransformers perform at least as well as post-hoc SAEs on sparse probing tasks, while out-performing on measures of feature absorption, steering effectiveness, and fine-grained causal interventions. Because subsequent computation acts only on features that survive the sparse bottleneck, the ParityTransformer's features are native to the model's forwards pass by construction, addressing the question of whether SAEs probe features the model actually uses during computation. We see this as a step toward training models whose internal representations are interpretable by design rather than recovered post hoc.
The results suggest that architectural routing mechanisms may have negligible impact on core semantic understanding, with representational divergence confined to extreme structural margins.
Rithin Nagaraj, Rupa Laalasa Oruganti, Prerna Subhashchandra Kunder et al.· 0 citations
OverRep is proposed, an Overcomplete Reparameterization framework for structured LLM pruning that temporarily overparameterizes the recovery module during training to absorb complex knowledge distilled from the original model.
This work proposes extracting mechanism mounts directly from linear sites by column-tiled SVD: each mount is a triple read as trigger, write, and strength, and evaluates mounts with a pre-registered suite judged on full-write energy lift rather than tile-local lift.
A systematic study of how pruning affects SAE behavior is presented and theoretically shows that, for a fixed SAE, its impact is governed by perturbation energy, a covariance-weighted norm.
Suchit Gupte, Xue-Ru Zhang, M. Khalili· 0 citations
A family of difference-informed pruning methods built upon this principle are introduced, suggesting that preserving output differences is a broadly useful and composable signal for post-training LLM sparsification.
Linghao Kong, Inimai Subramanian, Micah Adler et al.· 0 citations
This work replaces attention with a stack of autoencoder-based mixing modules, one operating over local neighborhoods, one over the full sequence, and one across attention heads, each compressing and reconstructing its input through a bottleneck, and its width is a hyperparameter rather than a training effect.
Narges Mokhtari, F. Haddadi, Ebrahim Rezaii· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.