Skip to content

CADSA: Content-Aware Dynamic Sparse Attention for Lightweight and Energy-Efficient Edge AI with Sub-Quadratic Complexity

Sep 2026 · Iraqi Journal for Computers and Informatics · 0 citations · 27 references
Advanced Neural Network Applications

TL;DR

CADSA is a content-aware dynamic sparse attention framework that lowers the quadratic complexity of self-attention to sub-quadratic for lightweight edge AI deployment, and fine-tuning strategy with gradient masking decreases the training memory overhead by 60%, allowing for practical on-device fine-tuning.

Abstract

This paper presents CADSA, a content-aware dynamic sparse attention framework that lowers the quadratic complexity of self-attention to sub-quadratic for lightweight edge AI deployment. A learnable token relevance scoring module dynamically prunes redundant attention paths by top-k selection in local neighborhoods, with complexity O(n·k·d). A variance-aware sparsity modulation mechanism is used for generative tasks, where the pruning ratio is adjusted according to the statistics of the local regions, focusing computational resources on detailed regions and simplifying uniform regions. Block-sparse execution kernels and bitmask compression reduce memory bandwidth by 32×, ensuring hardware efficiency. Experimental results show 2.1× inference speedup and 40% energy saving on Raspberry Pi 5, and 98.6% dense attention accuracy on ImageNet classification. On LSUN Bedrooms, CADSA achieves similar FID scores (12.7 vs. 12.3) and 2.4× faster sampling than diffusion models. The fine-tuning strategy with gradient masking decreases the training memory overhead by 60%, allowing for practical on-device fine-tuning.

Read PDF

Similar papers

Preprint Aug 2026

SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference

SPADE is presented, a training-free sparse-attention engine of three parts: a specification defining 3D blocking candidates and formalizing dynamic masks via Summarizer/Estimator expressions, and an executor with low-overhead index search, flash block-sparse attention, and kernel grouping.

Shang-Hao Liu, Renze Chen, Size Zheng et al. · 1 citation · ⚡1
Open access Aug 2026

A Collaborative Multi-Compression Acceleration Mechanism for Neural Networks in Keyword Spotting

This study proposes a collaborative multi-compression framework for lightweight deployment that combines adaptive importance-aware structured pruning, mixed-precision quantization, and quantization-aware multi-stage knowledge distillation on LiteKWS-Net.

Junbang Jiang, Rui Pu, Jin Li et al. · 0 citations
Preprint Aug 2026

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge

Edge LLM inference combines sparsity and low-bit quantization to meet device memory, latency, and power limits. Yet quantization shrinks weight payloads without proportionally reducing sparse metadata, so index traffic and nonzero extraction become critical SpMM bottlenecks. We introduce the Payload-to-Metadata Ratio (...

Tianhao Jiang, Hang Gu, Teng Wang et al. · 0 citations
Preprint Aug 2026

APT: Accelerating Diffusion Transformers via Attention Probability-Guided Pruning and Quantization

APT, a software-hardware co-designed accelerator for high-resolution DiTs that leverages attention probabilities as a unified importance metric to jointly optimize computation through fine-grained pruning and adaptive precision scaling, and is evaluated on SOTA DiT models including PixArt, Stable Diffusion 3, and FLUX.

Sungyeob Yoo, Seeyeon Kim, Joonyong Park et al. · 0 citations
Preprint Sep 2026

VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention

VC-Attention is proposed, a training-free low-bit attention framework that addresses diffusion Transformers and quantization scale by pairing Value smoothing with a fused probability Cast, and improves fidelity over low-bit baselines.

Xing-Yang Li, Dong-Yun Zou, Shi-Ning Zhang et al. · 1 citation · ⚡1
#artificial intelligence Preprint Aug 2026

Faster Than Flash: Exploiting Attention Sparsity for Efficient Long-Context Decoding

Faster Flash Decoding (FFD) is presented, a novel hardware-algorithm co-design framework designed to break the memory wall in long-context decoding and introduces the top-delta strategy, which dynamically filters blocks to achieve distribution-adaptive sparsity without global synchronization.

Zhigeng Liu, Zhiyuan Ning, Rui-Xiao Li et al. · 3 citations

Related blog posts

MIT News · Artificial Intelligence Sep 30, 2026

This game-playing AI is the new champ at Stratego

Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.

Microsoft Research Blog Sep 29, 2026

Introducing Quine: An AI research system designed for the complexity of biology

Biology doesn't operate in silos, and neither should the AI representation of it. Quine is an early-stage research effort to create a multimodal world model of biology. By connecting insights across biological scales and modalities, Quine helps scientists computationally search a space far larger than intuition allows and prioritize hypotheses before they reach the lab. Experimental results provide important feedback, helping researchers sharpen future research directions. The post Introducing Q…

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.