Skip to content

T-LoopFormer: Token-Level Elastic-Depth Looped Transformers for Latent Reasoning with Dynamic Routing

Sep 2026 · 1 citation · 49 references
Computer Science

TL;DR

This work proposes dynamic token-choice routing for looped transformers, enabling each token to adaptively determine its own number of loop iterations based on its hidden state, which can improve the token generation accuracy and validate the effectiveness of token-choice router and recursion-wise KV cache.

Abstract

Looped Transformers have recently demonstrated strong performance in both reasoning and language tasks by reusing a shared set of parameters across multiple iterations, achieving parameter efficiency without sacrificing representational power. Besides, looped Transformers perform inference directly in the latent space (latent reasoning) to reduce the number of tokens consumed during inference, thereby achieving improved sample efficiency. However, these models typically apply a fixed recursion depth uniformly to every token, leading to suboptimal compute allocation and leaving significant efficiency gains on the table. In this work, we propose dynamic token-choice routing for looped transformers, enabling each token to adaptively determine its own number of loop iterations based on its hidden state, which can improve the token generation accuracy. Moreover, we further introduce recursion-wise KV cache, which maintains an independent key-value cache for each recursion loop, this design ensures that tokens at different depths only attend to their corresponding cached states, effectively enabling faster autoregressive decoding. Extensive experiments show that T-LoopFormer achieves robust performance on language modeling and zero-shot reasoning tasks and our model can reach the lowest decoding latency, which validate the effectiveness of token-choice router and recursion-wise KV cache. Code is available at https://github.com/YuMingQian1234/T-LoopFormer

View source

Similar papers

#machine learning Preprint Sep 2026

Improving Test-Time Scaling with Adaptive Looped Transformers

Looped transformers have demonstrated promising parameter efficiency by reusing layers for latent computation. Prior studies compare looped and non-looped models at matched parameters or per-token FLOPs. However, to the best of our knowledge, whether looping improves test-time scaling as outputs grow longer remains und...

Yi-Chen You, Tianyu Fu, Ao-Song Feng et al. · 0 citations
#machine learning Preprint Sep 2026

LoopSpec: Pipelined Self-Speculative Decoding for Looped Transformers

Looped Transformers achieve strong performance with compact parameter sizes by repeatedly applying a shared stack of Transformer blocks across recurrent depths. However, they incur higher decoding latency than standard Transformer models of comparable parameter size because shared weights are accessed at every recurren...

SangLyul Cho, Lang-Qing Cui, Sehoon Kim et al. · 1 citation
Preprint Aug 2026

LoopMTP: A looped transformer guided by latent multi-token prediction

Multi-token prediction (MTP) supplies exactly the dense, forward-looking supervision the loop is missing, by softly aligning the hidden state of loop $t$ with the embedding of the token $t$ steps ahead, while a lightweight gate preserves useful information across iterations.

Behzad Shomali, Markus Frey, D. Berghaus et al. · 3 citations
#machine learning Preprint Sep 2026

Looped Transformers as Optimizers

Looped Transformers provide a parameter-efficient approach to depth scaling by repeatedly applying shared Transformer blocks. Recent reasoning models have likewise highlighted the value of scaling test-time computation through longer computation trajectories. However, the principles for designing effective loop transit...

Yu-Long Huang, Chen Jiang, Zhan-Peng Zhou et al. · 0 citations
#machine learning Preprint Sep 2026

FlashLoop: Fast and Memory-Efficient Looped Transformers via Lazy Updates

Looped Transformers have attracted substantial attention as a parameter-efficient approach to increasing computational depth through repeated application of shared Transformer blocks. However, their practical advantages over conventional Transformers remain under debate: each additional loop incurs another Transformer...

Wan-Qi Yang, Shi-Wei Liu · 0 citations
#machine learning Preprint Sep 2026

Trading Depth for Time in Recurrent Transformers

Recurrent Transformers increase computational depth through temporal recurrence, feeding each token's high-level hidden state into the computation of the next. This raises a natural question: is additional computation better spent on more temporal steps or greater physical depth? We investigate this question using Late...

Ze-Yi Huang, Xuehai He, Yong Jae Lee et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.