LiteTopK: Exploiting the Curse of Dimensionality for a Fused Indexer-TopK Kernel in Long-Context Sparse Attention
This study observes that sparse attention scores exhibit a score concentration phenomenon, where scores tend to fall within a narrow range, and proposes LITETOPK, an efficient fused Indexer-TopK kernel, which exploits the similarity of top-k candidate sets among neighboring tokens and proposes LITEDSA, which exploits the similarity of top-k candidate sets among neighboring tokens.