ELSAA: Efficient Low-Rank and Sparse Attention Approximation for Training Transformers
ELSAA is proposed, an efficient low-rank and sparse approximation of attention that gives a practical framework for constructing low-rank and sparse attention outputs without materializing the full quadratic score matrix, aiming to enable longer-context training while preserving both sharp token-level interactions and...