Skip to content

Author

Karen Mosoyan

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

A Controlled Study of Attention-Only Transformers

This work pretrain attention-only decoder transformers against standard transformers matched separately for parameter count, training FLOPs, and depth (2 to 48 layers), for up to 105B tokens at 6M to 87M parameters, and localizes the remaining gap to parametric recall.

Henry Ndubuaku, Karen Mosoyan, Jakub Mroz et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.