Skip to content

ELSAA: Efficient Low-Rank and Sparse Attention Approximation for Training Transformers

Jul 2026 · arXiv.org · Vol abs/2607.20214 · 0 citations · 44 references
Computer Science

TL;DR

ELSAA is proposed, an efficient low-rank and sparse approximation of attention that gives a practical framework for constructing low-rank and sparse attention outputs without materializing the full quadratic score matrix, aiming to enable longer-context training while preserving both sharp token-level interactions and broad contextual mixing.

Abstract

The quadratic $N\times N$ attention score matrix remains a central obstacle to extending Transformers to longer input lengths. Existing efficient attention methods usually reduce this bottleneck by either imposing sparsity, so that each query attends to only a small subset of keys, or by using low-rank/kernel sketches, so that global interactions are compressed into a lower-dimensional representation. We propose \emph{ELSAA}, an efficient low-rank and sparse approximation of attention. Importantly, ELSAA does \emph{not} decompose the learned projection or output matrices of the Transformer into sparse and low-rank factors. Instead, after dense projections produce $Q,K,V$, ELSAA approximates the induced attention score operator itself: a sparse branch captures selected high-similarity interactions, while a low-rank branch summarizes diffuse global interactions. Since the two branches can be normalized over supports with very different denominator mass, ELSAA introduces a denominator-aware fusion term that scales the sparse branch according to its estimated attention mass relative to the low-rank branch. This gives a practical framework for constructing low-rank and sparse attention outputs without materializing the full quadratic score matrix, aiming to enable longer-context training while preserving both sharp token-level interactions and broad contextual mixing.

View source

Similar papers

#machine learning Preprint Aug 2026

FLARE++: Low-rank attention with dynamic attention routing

FLARE++ is a low-rank attention architecture with input-conditioned routing queries that reduces FLARE's error by $25\% on average across five standard PDE benchmarks, achieving the lowest errors among the efficient models compared.

Vedant Puri, Y. Zhang, Levent Burak Kara · 0 citations
Preprint Sep 2026

RoLA: Rotary-Positioned Low-Rank Linear Attention for Efficient Diffusion Transformers

Diffusion Transformers (DiTs) achieve strong video generation quality, but their dense spatiotemporal self-attention scales quadratically with sequence length and quickly becomes the dominant inference bottleneck. Sparse low-rank hybrids alleviate this cost by combining a local sparse branch with a global compressed br...

Ze-Kun Zhang, Yi-Xiang Cai, Yu-Xi Liu et al. · 1 citation
#artificial intelligence Preprint Aug 2026

Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models

This work introduces SparsePR, which combines Response-Coupled Partitioning with Probe-Fitted Residual Reconstruction and consistently reduces attention-reconstruction error across four heterogeneous video generation and world models.

P. Taghavi, Reza Langari, Gaurav Pandey · 1 citation
Open access Aug 2026

Resource-Efficient Pruning for Transformer via Low-Rank Importance Estimation

REP-LIE leverages the gradients of LoRA low-rank matrices to estimate the importance of weights without requiring full gradient computation, and a stability score is introduced, serving as the basis for iterative pruning of unimportant model parameters.

Peng Liu, Hui-Bing Zeng, Yi-Qun Zhang et al. · 0 citations
Preprint Aug 2026

Fine-Tuning of Transformer models with Frames

The experiments show that FrameFT achieves performance on par with/exceeding state-of-the-art PEFT techniques, but needs far fewer trainable parameters.

Harshavardhan Adepu, Li Zhang, Sanjiv Kumar et al. · 0 citations
#machine learning Preprint Sep 2026

PQ-HSA: Reusing Product-Quantized Scores for Hybrid Sparse-Approximate Attention

PQ-HSA (hybrid sparse-approximate attention) attends the selected tokens with their original keys and values, and the unselected tokens, the background, enter the same softmax through those scores, summed per inverted list and multiplied by the list's mean value.

Kun-Ming Shao, Jie-Run Chen, Yan-Li Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.