Skip to content
Review Open access

A Survey of Dynamic Token Computation in Transformers: Taxonomy, Stability, and Budget-Aware Evaluation

2026 · IEEE Access · Vol 14, pp. 119498-119537 · 0 citations · 164 references
Computer Science

TL;DR

This survey contributes a structured vocabulary, a budget- and stability-aware evaluation framework, and a matched benchmarking protocol for analyzing whether dynamic token methods provide reliable and deployment-relevant efficiency gains.

Abstract

Transformers have delivered strong performance across vision, language and multimodal tasks but their computational cost grows rapidly with token count, creating a major obstacle for real-time inference, edge deployment and other resource-constrained settings. A central challenge is that many input tokens contribute unequally to final predictions, yet conventional Transformer pipelines still process them uniformly, resulting in redundant computation, latency overhead and inefficient resource use. To address this problem, this survey reviews dynamic token computation in Transformers, an important research area that seeks to reduce unnecessary computation while preserving predictive quality. The survey includes token pruning and dropping, token merging and aggregation, conditional token routing and adaptive depth or early-exit strategies across Transformer-family architectures. Its objective is to provide a taxonomy-driven evaluation framework that unifies existing methods under common token-decision axes and assesses them through routing stability, budget reliability, overhead-aware latency and matched benchmark criteria. The synthesis shows that prior work can be grouped by operation type, decision design, training strategy and budget control. Although reported efficiency gains are often promising, they depend strongly on selection overheads, hardware settings and whether theoretical compute reduction translates into actual latency improvement. Recurring weaknesses include inconsistent evaluation protocols, limited reproducibility, unstable routing decisions and weak guarantees under perturbation or distribution shift. This survey contributes a structured vocabulary, a budget- and stability-aware evaluation framework, and a matched benchmarking protocol for analyzing whether dynamic token methods provide reliable and deployment-relevant efficiency gains.

Read PDF

Similar papers

#artificial intelligence Preprint Aug 2026

Pro-Router: Token-Aware Progressive Model Routing with Adaptive Edge-Cloud Collaboration for Efficient Multimodal LLM Inference

Pro-Router is proposed, a token-aware progressive model routing method with adaptive edge-cloud collaboration for efficient multimodal LLM inference and achieves the highest routing accuracy and improves routing speed by more than 10x.

Xin-Yuan Gui, Shao-Wen Wang, Sheng Sun et al. · 0 citations
Preprint Aug 2026

Sparse Token Routing in Efficient Transformers

Efficient-transformer research often motivates token pruning and adaptive computation with the claim that not all tokens require equal computational effort. We test this claim end to end using SEWN, a two-stream Transformer that routes tokens through either lightweight or full-capacity processing using a learned gate....

Sai Krishna Arthanari, JaeHyeong Chang, Chengzhe Sun et al. · 0 citations
Preprint Sep 2026

SafeDepth: Safety-Aware Token-Level Adaptive Computation

Recent studies suggest that not every token needs to pass through all Transformer layers, motivating token-level adaptive models that selectively skip layers to reduce computation. Our experiments show that these execution choices also affect safety: existing token-level adaptive reasoning models exhibit higher harmful...

Ni-Zhang Li, Ian G. Harris · 0 citations
#artificial intelligence Preprint Sep 2026

T-LoopFormer: Token-Level Elastic-Depth Looped Transformers for Latent Reasoning with Dynamic Routing

This work proposes dynamic token-choice routing for looped transformers, enabling each token to adaptively determine its own number of loop iterations based on its hidden state, which can improve the token generation accuracy and validate the effectiveness of token-choice router and recursion-wise KV cache.

Ming-Qian Yu, Wen-Peng Zhang, Shao-Bo Cui et al. · 2 citations
#machine learning Preprint Aug 2026

LoGo: Token-Level Dynamic Local-Global Attention

LoGo, a token-level dynamic local-global attention mechanism that uses attention span as a direct proxy for attention budget allocation, is proposed and results suggest that learned token-level span allocation is an effective and scalable way to improve the long-context performance-compute trade-off.

Yu-Qi Pan, Zheng Li, Bo-Hao Tang et al. · 1 citation
#artificial intelligence Preprint Sep 2026

On the Token Value Inequality in Efficient Reasoning

This work presents a diagnostic and optimization framework grounded in a key empirical finding: the value of tokens within a CoT reasoning sequence is highly non-uniform, and this non-uniformity can be effectively characterized by token-level log probability signals.

Run-Jia Zeng, Hang Hua, Yiyang Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.