Skip to content

RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models

Sep 2026 · 0 citations · 26 references
Computer Science

TL;DR

RBS-Attention is proposed, a training-free sparse-prefill method with two complementary selection branches that controls the contribution of rescue blocks while preserving regular block-sparse FlashAttention execution and supports radius-adaptive dual-branch selection as an effective approach to long-context prefill.

Abstract

Long-context large language model inference is increasingly limited by prefill, where dense self-attention processes the entire prompt before generation begins. Sparse block selection can reduce this cost, but a block centroid may hide a highly relevant token among many irrelevant ones. We call this failure mode mean dilution and propose RBS-Attention, a training-free sparse-prefill method with two complementary selection branches. A centroid base branch captures average relevance, while a rescue branch uses the maximum key-block radius and its prompt-, layer-, and head-dependent distribution to identify blocks at risk of underestimation. Independently thresholding the two branches and combining their masks controls the contribution of rescue blocks while preserving regular block-sparse FlashAttention execution. On H100 GPUs, RBS-Attention achieves 20.65$\times$ standalone prefill-attention speedup, 11.92$\times$ vLLM prefill-attention speedup, and 5.97$\times$ end-to-end time-to-first-token speedup at 128K on Qwen3-30B-A3B-Instruct-2507-FP8. On the dense Qwen3-32B model, it obtains 88.65 overall RULER accuracy versus 89.52 for dense attention; LongBench-v2, InfiniteBench, and Video-MME provide additional quality evaluation. Supporting experiments measure actual retention, compare selectors at matched density, and characterize block-size, threshold, and memory behavior. Together, these results support radius-adaptive dual-branch selection as an effective approach to long-context prefill.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Context Poisoning as Extreme-Value Attention Interference in Long-Context Language Models

Large language models can process increasingly long prompts, yet their ability to locate and use decisive evidence may degrade as irrelevant or confusable context is added. We formulate this phenomenon, which we call context poisoning, as extreme-value interference in attention: the decisive-evidence score is upper-bou...

Meysam Ghaffari, Nina Fatehi, Bhaskar Sen et al. · 0 citations
#machine learning Preprint Aug 2026

LoGo: Token-Level Dynamic Local-Global Attention

LoGo, a token-level dynamic local-global attention mechanism that uses attention span as a direct proxy for attention budget allocation, is proposed and results suggest that learned token-level span allocation is an effective and scalable way to improve the long-context performance-compute trade-off.

Yu-Qi Pan, Zheng Li, Bo-Hao Tang et al. · 1 citation
#machine learning Preprint Sep 2026

Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding

This work introduces Elastic Threshold Attention, an end-to-end trainable architecture that rivals dense model quality under hardware-aligned block-sparse decoding, and introduces an offline calibration algorithm for domain-specific deployments that freezes per-head constant thresholds, cutting attention compute by an...

Themistoklis Haris, Henry Li, Maryam Karimzadehgan · 1 citation · ⚡1
#natural language process... Preprint Sep 2026

Scaling Parameter and Context in Attention: Native Sparse Attention from Mixture-of-Head

NAMOH, an architecture-native sparse attention mechanism that activates only its assigned tokens and performs causal attention within this subsequence, is introduced, and it is hoped this work offers a new path for scaling attention, with parameter scaling directly enabling context scaling.

Zi-Zhuo Fu, Run-Sheng Wang, Meng Li · 0 citations
#natural language process... Preprint Sep 2026

CEDAR: Error-Bounded Residual Routing for Efficient Long-Context Attention

This work introduces Coarse-to-fine Error-aware Dynamic Attention Routing (CEDAR), a coarse-to-fine method that keeps the language model frozen while preserving global coverage and derives an output-error bound governed by within-chunk key/value dispersion and uses it to allocate a variable refinement budget.

Si-Yu Li, Dong Wang, Jie Zhou et al. · 0 citations
#natural language process... Preprint Oct 2026

Role-aware Heuristic Episodic Attention for Conversational LLMs

Large language models often lose track of persistent instructions and relevant information as multi-turn conversations grow. We study this cumulative contextual decay through three related failure modes: attention pollution, dilution, and drift. We propose REA (Role-aware Heuristic Episodic Attention), a context-manage...

Wan-Yang Hong, Zhao-Ning Zhang, Yi Chen et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.