Skip to content

A Controlled Study of Attention-Only Transformers

Jul 2026 · arXiv.org · Vol abs/2607.18363 · 1 citation · 27 references
Computer Science

TL;DR

This work pretrain attention-only decoder transformers against standard transformers matched separately for parameter count, training FLOPs, and depth (2 to 48 layers), for up to 105B tokens at 6M to 87M parameters, and localizes the remaining gap to parametric recall.

Abstract

Feed-forward networks hold two thirds of a transformer's non-embedding parameters, yet the architecture has not received a necessity test that controls parameters, compute, and depth at once. We pretrain attention-only decoder transformers (Simple Attention Networks, SANs) against standard transformers matched separately for parameter count, training FLOPs, and depth (2 to 48 layers), for up to 105B tokens at 6M to 87M parameters. Deleting feed-forward layers in place is costly: the standard transformer leads by 0.47 nats at matched depth and 0.26 nats at matched FLOPs. Reallocating the freed budget into attention depth closes the gap: at matched parameters the difference is 0.006 nats (0.27 percent of loss), reproducible to one part in ten thousand across seed pairs, shrinking across 5B, 30B, and 105B budgets, and holding near 0.02 nats across a 29x size range. Three measurements localize the remaining gap to parametric recall: attention-only models are better on context-grounded answers and worse where knowledge must come from weights. Weight spectra show why: routing matrices (Q/K) crystallize early, content matrices accumulate rank slowly, and removing feed-forward layers relocates this accumulation to the attention output projection. QK-normalization, not feed-forward layers or residual gating, keeps 48-layer attention-only stacks trainable. The deficit concentrates on low-context query prediction and localizes there entirely by the largest budget. A pre-registered test confirms the account: it predicts a 0.02 to 0.05 nat gap on knowledge-dense web text; a matched pair trained on fineweb-edu measures 0.040. Within the tested regime, attention does the rest.

View source

Similar papers

Preprint Aug 2026

TANGO: Treating Tokens as Operators

Transformers separate cross-token mixing in self-attention from token-wise transformation in feed-forward networks. We ask whether combining these operations can lower predictive loss under fixed data and parameter budgets. To do so, we introduce the Token-Aggregated Nonlinear Gating Operator (TANGO) model. TANGO compu...

Joshua Nunley · 0 citations
#artificial intelligence Preprint Sep 2026

From Attention Sensitivity to Layer Role: Revisiting Mixed-Precision Quantization of Transformers

Most post-training quantization pipelines fit each weight matrix to its pretrained counterpart, one matrix at a time. Whether that proxy tracks what an attention block actually computes, or how errors in the Q, K and V projections compound inside the softmax, is rarely checked. We write the objective on the attention o...

Nafiseh HosseinpourFardi, Negar Alihadi, Mahmoudreza Babaei et al. · 0 citations
Preprint Aug 2026

Full-bandwidth transformer

This work trains 1B-parameter full-bandwidth transformers on up to 400B tokens and finds that latent feedback improves validation loss, 5-shot language-model evaluation, math and coding generation, and instruction-tuned performance.

Xi Wang, Ziyang Cai, Zheng Zhan et al. · 5 citations
#artificial intelligence Preprint Sep 2026

What Does Layer-Importance Reveal About Transformers and State-Space Models?

Transformers and state-space models (SSMs) are the two dominant families of sequence models, and a central open question is how far the analytical knowledge built for transformers transfers to SSMs. We address this through the lens of layer importance which underpins compression, selective fine-tuning, and interpretabi...

Istabrak Abbes, Nizar Islah, Irina Rish et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Pretrained transformers use little of their depth to follow references in context. Thirteen base models reliably follow only 1.4-3.6 lines, and extra pretrained loops add little. A task-trained rank-8 LoRA at one early layer extends this computation with all model weights frozen. Qwen3-8B improves from 15.5% to 99% exa...

Ze-Hao Jin, Rui-Xuan Deng, Jun-Ran Wang · 0 citations
#artificial intelligence Preprint Sep 2026

How to Loop MoE: Flatten the Experts, Untie the Attention

Looped Transformers reuse one block of layers several times: by spending extra computation they push a model of fixed size further, and so use its parameters more fully; while sparse mixture-of-experts (MoE) models activate only a few of many experts for each token. Looped MoE bridges these two design philosophies and...

Shouren Wang, Chuan Ma, Mohsen Hariri et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.