Preprint
Jul 2026
FastTPS: An Optimized Method for LLM Token Phase for AI accelerators
This paper presents FastTPS, a high performance and low-precision loss method for accelerating the token-phase in LLM inference on general AI accelerators which includes three key components: AI accelerator-enabled reloading-free KV Cache concatenation which decreases memory access overhead as well as enables full fusion of Attention.
Wenzong Yang, Danyang Zhang, Kunteng Cao et al.
· 0 citations