RBS-Attention is proposed, a training-free sparse-prefill method with two complementary selection branches that controls the contribution of rescue blocks while preserving regular block-sparse FlashAttention execution and supports radius-adaptive dual-branch selection as an effective approach to long-context prefill.
Abstract
Long-context large language model inference is increasingly limited by prefill, where dense self-attention processes the entire prompt before generation begins. Sparse block selection can reduce this cost, but a block centroid may hide a highly relevant token among many irrelevant ones. We call this failure mode mean dilution and propose RBS-Attention, a training-free sparse-prefill method with two complementary selection branches. A centroid base branch captures average relevance, while a rescue branch uses the maximum key-block radius and its prompt-, layer-, and head-dependent distribution to identify blocks at risk of underestimation. Independently thresholding the two branches and combining their masks controls the contribution of rescue blocks while preserving regular block-sparse FlashAttention execution. On H100 GPUs, RBS-Attention achieves 20.65$\times$ standalone prefill-attention speedup, 11.92$\times$ vLLM prefill-attention speedup, and 5.97$\times$ end-to-end time-to-first-token speedup at 128K on Qwen3-30B-A3B-Instruct-2507-FP8. On the dense Qwen3-32B model, it obtains 88.65 overall RULER accuracy versus 89.52 for dense attention; LongBench-v2, InfiniteBench, and Video-MME provide additional quality evaluation. Supporting experiments measure actual retention, compare selectors at matched density, and characterize block-size, threshold, and memory behavior. Together, these results support radius-adaptive dual-branch selection as an effective approach to long-context prefill.
Large language models can process increasingly long prompts, yet their ability to locate and use decisive evidence may degrade as irrelevant or confusable context is added. We formulate this phenomenon, which we call context poisoning, as extreme-value interference in attention: the decisive-evidence score is upper-bou...
Meysam Ghaffari, Nina Fatehi, Bhaskar Sen et al.· 0 citations
LoGo, a token-level dynamic local-global attention mechanism that uses attention span as a direct proxy for attention budget allocation, is proposed and results suggest that learned token-level span allocation is an effective and scalable way to improve the long-context performance-compute trade-off.
Yu-Qi Pan, Zheng Li, Bo-Hao Tang et al.· 1 citation
This work introduces Elastic Threshold Attention, an end-to-end trainable architecture that rivals dense model quality under hardware-aligned block-sparse decoding, and introduces an offline calibration algorithm for domain-specific deployments that freezes per-head constant thresholds, cutting attention compute by an...
Themistoklis Haris, Henry Li, Maryam Karimzadehgan· 1 citation· ⚡1
NAMOH, an architecture-native sparse attention mechanism that activates only its assigned tokens and performs causal attention within this subsequence, is introduced, and it is hoped this work offers a new path for scaling attention, with parameter scaling directly enabling context scaling.
This work introduces Coarse-to-fine Error-aware Dynamic Attention Routing (CEDAR), a coarse-to-fine method that keeps the language model frozen while preserving global coverage and derives an output-error bound governed by within-chunk key/value dispersion and uses it to allocate a variable refinement budget.
Large language models often lose track of persistent instructions and relevant information as multi-turn conversations grow. We study this cumulative contextual decay through three related failure modes: attention pollution, dilution, and drift. We propose REA (Role-aware Heuristic Episodic Attention), a context-manage...
Wan-Yang Hong, Zhao-Ning Zhang, Yi Chen et al.· 0 citations
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.