Preprint
Jul 2026
RIS-Kernel: A Model-Agnostic Architecture for Long-Context LLM Inference via Sparse Attention
In controlled evaluations at 32,768 tokens, RIS-Stochastic at 1% density and 70 ensemble seeds achieves 75.00% accuracy, outperforming the native dense baseline, demonstrating that sparse attention acts as a regularizer: low density over multiple seeds filters out sequence-level noise, whereas higher density reintroduces distractor noise.
A. R. Santos
· 1 citation