Jul 2026
A Controlled Study of Attention-Only Transformers
This work pretrain attention-only decoder transformers against standard transformers matched separately for parameter count, training FLOPs, and depth (2 to 48 layers), for up to 105B tokens at 6M to 87M parameters, and localizes the remaining gap to parametric recall.
Henry Ndubuaku, Karen Mosoyan, Jakub Mroz et al.
· arXiv.org · 1 citation