Adaptive Model Compression (AMC), a saliency-driven framework that dynamically allocates hardware resources based on token importance, effectively extends the battery life of mobile devices by utilizing high-definition compute only where necessary, maintaining robust performance with a marginal 3.6% accuracy trade-off.
Abstract
Deploying large-scale transformer models on resource-constrained edge devices remains a challenge due to the high energy and memory overhead inherent in static inference, which processes simple and complex tokens with uniform intensity. To address this, we propose Adaptive Model Compression (AMC), a saliency-driven framework that dynamically allocates hardware resources based on token importance. By implementing a multi-tier architecture, our system identifies critical high-saliency information for full-precision processing while aggressively reducing the rank and bit-width of less significant data. Experimental results demonstrate that AMC achieves a 59.2% reduction in system energy and a 2.24x increase in throughput on 45nm CMOS hardware. This approach effectively extends the battery life of mobile devices by utilizing high-definition compute only where necessary, maintaining robust performance with a marginal 3.6% accuracy trade-off.
CADSA is a content-aware dynamic sparse attention framework that lowers the quadratic complexity of self-attention to sub-quadratic for lightweight edge AI deployment, and fine-tuning strategy with gradient masking decreases the training memory overhead by 60%, allowing for practical on-device fine-tuning.
Abdullah Ghanim Jaber, Abeer Ahmed Ali, Ibrahim Alshammari et al.· Iraqi Journal for Computers...· 0 citations
DeVIT, an acceleration method for vision transformers that leverages differential computation to enable multiplier-less matrix multiplication and introduces another useful property: value locality.
Reyhaneh Hosseinzadeh, Parham Zilouchian Moghaddam, Mehdi Modarressi· 0 citations
APT, a software-hardware co-designed accelerator for high-resolution DiTs that leverages attention probabilities as a unified importance metric to jointly optimize computation through fine-grained pruning and adaptive precision scaling, and is evaluated on SOTA DiT models including PixArt, Stable Diffusion 3, and FLUX.
Sungyeob Yoo, Seeyeon Kim, Joonyong Park et al.· 0 citations
A reparameterized transformer framework that integrates High-Rank Factorization (HRF) during training, layer merging at inference, and dynamic, load-balanced distributed inference across multiple devices is proposed, highlighting that reparameterized transformers, coupled with adaptive distributed inference and ultra-l...
H. Esmaeili, M. A. Afsharkazemi, R. Radfar et al.· Turkish Journal of Electrica...· 0 citations
Findings confirm that combining complementary compression strategies yields substantially better performance-efficiency trade-offs than any single technique applied in isolation.
Upma Sharma Archana· International Journal of Res...· 0 citations
A compact and configurable event-driven autoencoder that efficiently compresses neuromorphic data while preserving essential spatiotemporal structure for downstream inference and demonstrates the potential of compact event-driven models to advance environmentally conscious, low-power AI systems for high-speed perceptio...
Riadul Islam, Joey Mulé, Dhandeep Challagundla et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.