Skip to content
Preprint

Full-bandwidth transformer

Aug 2026 · 5 citations · 53 references
Computer Science

TL;DR

This work trains 1B-parameter full-bandwidth transformers on up to 400B tokens and finds that latent feedback improves validation loss, 5-shot language-model evaluation, math and coding generation, and instruction-tuned performance.

Abstract

Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through model depth. Dense attention gives each token broad horizontal access to the past, but the vertical feedback channel between decoding steps remains narrow: only the sampled token returns to the bottom of the stack, while the top-layer hidden state is discarded. We introduce the full-bandwidth transformer, which widens this channel with latent feedback: at each decoding step, the previous top-layer hidden state is fused with the sampled token embedding through a gated linear unit and fed back as the next input. Latent feedback lets non-verbalized computation re-enter the stack with a renewed depth budget, while preserving the standard transformer architecture, KV cache, and language-modeling objective. To train full-bandwidth transformers without losing parallel teacher forcing, we use a scheduled multi-pass objective that introduces latent feedback late in pretraining and mixes a small fraction of deeper feedback passes for stability. We train 1B-parameter full-bandwidth transformers on up to 400B tokens and find that latent feedback improves validation loss, 5-shot language-model evaluation, math and coding generation, and instruction-tuned performance. With negligible per-token decoding overhead, full-bandwidth transformers match or approach standard transformers trained with roughly 1.5x more tokens, and manage to produce shorter reasoning when no off-policy templates are provided.

View source

Similar papers

#machine learning Preprint Sep 2026

Trading Depth for Time in Recurrent Transformers

Recurrent Transformers increase computational depth through temporal recurrence, feeding each token's high-level hidden state into the computation of the next. This raises a natural question: is additional computation better spent on more temporal steps or greater physical depth? We investigate this question using Late...

Ze-Yi Huang, Xuehai He, Yong Jae Lee et al. · 0 citations
Preprint Aug 2026

TANGO: Treating Tokens as Operators

Transformers separate cross-token mixing in self-attention from token-wise transformation in feed-forward networks. We ask whether combining these operations can lower predictive loss under fixed data and parameter budgets. To do so, we introduce the Token-Aggregated Nonlinear Gating Operator (TANGO) model. TANGO compu...

Joshua Nunley · 0 citations
#natural language process... Preprint Sep 2026

Pretraining Latent Information Feedback Transformers with Teacher Supervision

Transformer language models (LMs) are feed-forward: deep-layer representations are never fed back to shallower layers, and the only pathway for information to flow downward across generation steps is the decoded token. This narrow channel forces models to recompute intermediate results and to discard alternative contin...

Dor Tirosh, Ido Amos, Mor Geva · 0 citations
#small language model Preprint Sep 2026

Compressing Streaming Neural Audio Encoders via Latent-Space Distillation

This work studies how to compress a tokenizer by distillation, taking as the supervision target neither the discrete token nor the output distribution but the pre-quantizer latent the model actually consumes - the last representation the two token interfaces share.

Prasanth Yadla, Mohammad Samragh, Dongseong Hwang et al. · 0 citations
#machine learning Preprint Aug 2026

WhiteMatter: All-to-All Cross-Layer Connections via KV Source Mixing

When generating text, a Transformer produces representations of past tokens at every layer, but each layer can normally use only representations from the same depth. This restriction prevents the model from fully reusing information it has already computed. We introduce WhiteMatter, which allows every layer to draw on...

Wen-Bo Zhang, Xiang Ren · 0 citations
#artificial intelligence Preprint Sep 2026

From Attention Sensitivity to Layer Role: Revisiting Mixed-Precision Quantization of Transformers

Most post-training quantization pipelines fit each weight matrix to its pretrained counterpart, one matrix at a time. Whether that proxy tracks what an attention block actually computes, or how errors in the Q, K and V projections compound inside the softmax, is rarely checked. We write the objective on the attention o...

Nafiseh HosseinpourFardi, Negar Alihadi, Mahmoudreza Babaei et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.