Skip to content

DeepLoop: Depth Scaling for Looped Transformers

Jul 2026 · arXiv.org · Vol abs/2607.13491 · 2 citations · 39 references
Computer Science

TL;DR

The results show that stable recurrent depth requires residual scaling rules that account for parameter visits, not only nominal layer count, and DeepLoop is neutral when no physical block is revisited and improves validation loss and downstream accuracy once recurrent depth is activated.

Abstract

Looped Transformers scale sequential computation by applying a compact stack of physical blocks for multiple rounds, increasing unrolled depth without increasing stored parameters. This reuse changes the residual-scaling problem: in an untied Transformer, each residual branch receives and applies its own parameter update, whereas in a looped Transformer one shared update aggregates gradients from repeated visits and is read back by those same visits in the next linearized forward pass. We formalize this tied-depth effect through a first-order perturbation bound controlled by a visit-alignment coefficient $\kappa_R$. The bound recovers the DeepNorm exponent when visits decorrelate, but in the conservative aligned regime it requires the exponent to increase from $1/4$ to $1/2$ as loop count grows at fixed physical depth. The resulting method, \textbf{DeepLoop}, keeps the Post-LN DeepNorm architecture and sets $\alpha=(2N)^{1/2}$ and $\beta=(8N)^{-1/2}$ for unrolled depth $N$. On GPT-style looped language models at GPT-2 small and GPT-2 medium scale, DeepLoop is neutral when no physical block is revisited and improves validation loss and downstream accuracy once recurrent depth is activated. These results show that stable recurrent depth requires residual scaling rules that account for parameter visits, not only nominal layer count.

View source

Similar papers

#machine learning Preprint Sep 2026

FlashLoop: Fast and Memory-Efficient Looped Transformers via Lazy Updates

Looped Transformers have attracted substantial attention as a parameter-efficient approach to increasing computational depth through repeated application of shared Transformer blocks. However, their practical advantages over conventional Transformers remain under debate: each additional loop incurs another Transformer...

Wan-Qi Yang, Shi-Wei Liu · 0 citations
#machine learning Preprint Sep 2026

LoopSpec: Pipelined Self-Speculative Decoding for Looped Transformers

Looped Transformers achieve strong performance with compact parameter sizes by repeatedly applying a shared stack of Transformer blocks across recurrent depths. However, they incur higher decoding latency than standard Transformer models of comparable parameter size because shared weights are accessed at every recurren...

SangLyul Cho, Lang-Qing Cui, Sehoon Kim et al. · 1 citation
#artificial intelligence Preprint Sep 2026

T-LoopFormer: Token-Level Elastic-Depth Looped Transformers for Latent Reasoning with Dynamic Routing

This work proposes dynamic token-choice routing for looped transformers, enabling each token to adaptively determine its own number of loop iterations based on its hidden state, which can improve the token generation accuracy and validate the effectiveness of token-choice router and recursion-wise KV cache.

Ming-Qian Yu, Wen-Peng Zhang, Shao-Bo Cui et al. · 1 citation
Preprint Sep 2026

Depth through recurrence: Looped transformers for flow-matching TTS

We study how to organize Transformer depth through recurrence in flow-matching text-to-speech, varying the amount, order, and place?ment of weight reuse. Seven layouts perform 18 block calls per network evaluation under a common training objective and sam?pler. On Seed-TTS and LibriSpeech-PC, SEQUENCE applies each of n...

Jia-Bao Ai, Peng Han, Yu-Chen Song et al. · 0 citations
#machine learning Preprint Aug 2026

Universal Transformers for Circuit Computations: Perfect Length Generalization in Tiny Transformers

Though the construction provably evaluates Boolean expressions -- a universal symbolic computation -- of arbitrary length perfectly, in other experiments it is demonstrated that the transformer variant can learn and generalize perfectly on other common length generalization benchmarks, including modular arithmetic and...

Takuya Ito, Ruchir Puri, Murray Campbell et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.