Jul 2026
Scaling Interpretable Transformers with Parity Bottleneck Layers
The ParityTransformer is introduced, a GPT-2-scale architecture whose intermediate representations are efficient and wide / sparse by design, and seen as a step toward training models whose internal representations are interpretable by design rather than recovered post hoc.
Andrew Mack, K. Tou, Mark C. Henry et al.
· arXiv.org · 0 citations