Conference
Open access
2026
RoBSA: RoPE-based Blockwise Sparse Multi-head Latent Attention
This paper introduces RoPE-based Block-wise Sparse Attention (RoBSA), a method designed specifically for MLA during the decoding stage of model inference that significantly reduces end-to-end inference latency in the decoding stage by up to 2 .
Xinyu Shi, Kairong Luo, Zhen Zheng et al.
· Annual Meeting of the Associ... · 0 citations