Preprint
Aug 2026
FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving
This paper introduces a mean correction term that effectively suppresses the approximation error, keeping performance degradation manageable even at extreme sparsity levels, and redesigns the sparse attention operator with PackGQA memory access, warp specialization, and pingpong pipelining.
Qihang Fan, Huaibo Huang, Zhiying Wu et al.
· 0 citations