Sep 2026· Iraqi Journal for Computers and Informatics· 0 citations· 27 references
Advanced Neural Network Applications
TL;DR
CADSA is a content-aware dynamic sparse attention framework that lowers the quadratic complexity of self-attention to sub-quadratic for lightweight edge AI deployment, and fine-tuning strategy with gradient masking decreases the training memory overhead by 60%, allowing for practical on-device fine-tuning.
Abstract
This paper presents CADSA, a content-aware dynamic sparse attention framework that lowers the quadratic complexity of self-attention to sub-quadratic for lightweight edge AI deployment. A learnable token relevance scoring module dynamically prunes redundant attention paths by top-k selection in local neighborhoods, with complexity O(n·k·d). A variance-aware sparsity modulation mechanism is used for generative tasks, where the pruning ratio is adjusted according to the statistics of the local regions, focusing computational resources on detailed regions and simplifying uniform regions. Block-sparse execution kernels and bitmask compression reduce memory bandwidth by 32×, ensuring hardware efficiency. Experimental results show 2.1× inference speedup and 40% energy saving on Raspberry Pi 5, and 98.6% dense attention accuracy on ImageNet classification. On LSUN Bedrooms, CADSA achieves similar FID scores (12.7 vs. 12.3) and 2.4× faster sampling than diffusion models. The fine-tuning strategy with gradient masking decreases the training memory overhead by 60%, allowing for practical on-device fine-tuning.
SPADE is presented, a training-free sparse-attention engine of three parts: a specification defining 3D blocking candidates and formalizing dynamic masks via Summarizer/Estimator expressions, and an executor with low-overhead index search, flash block-sparse attention, and kernel grouping.
This study proposes a collaborative multi-compression framework for lightweight deployment that combines adaptive importance-aware structured pruning, mixed-precision quantization, and quantization-aware multi-stage knowledge distillation on LiteKWS-Net.
Junbang Jiang, Rui Pu, Jin Li et al.· Symmetry· 0 citations
Edge LLM inference combines sparsity and low-bit quantization to meet device memory, latency, and power limits. Yet quantization shrinks weight payloads without proportionally reducing sparse metadata, so index traffic and nonzero extraction become critical SpMM bottlenecks. We introduce the Payload-to-Metadata Ratio (...
Tianhao Jiang, Hang Gu, Teng Wang et al.· 0 citations
APT, a software-hardware co-designed accelerator for high-resolution DiTs that leverages attention probabilities as a unified importance metric to jointly optimize computation through fine-grained pruning and adaptive precision scaling, and is evaluated on SOTA DiT models including PixArt, Stable Diffusion 3, and FLUX.
Sungyeob Yoo, Seeyeon Kim, Joonyong Park et al.· 0 citations
VC-Attention is proposed, a training-free low-bit attention framework that addresses diffusion Transformers and quantization scale by pairing Value smoothing with a fused probability Cast, and improves fidelity over low-bit baselines.
Faster Flash Decoding (FFD) is presented, a novel hardware-algorithm co-design framework designed to break the memory wall in long-context decoding and introduces the top-delta strategy, which dynamically filters blocks to achieve distribution-adaptive sparsity without global synchronization.
Zhigeng Liu, Zhiyuan Ning, Rui-Xiao Li et al.· 3 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 1, 2026
Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.
Biology doesn't operate in silos, and neither should the AI representation of it. Quine is an early-stage research effort to create a multimodal world model of biology. By connecting insights across biological scales and modalities, Quine helps scientists computationally search a space far larger than intuition allows and prioritize hypotheses before they reach the lab. Experimental results provide important feedback, helping researchers sharpen future research directions. The post Introducing Q…
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.