The MiCo framework is proposed, a holistic MPQ exploration and deployment framework for edge AI applications that adopts a novel optimization algorithm to search for accuracy-optimal quantization configurations under strict latency constraints and is extended to MiCoPro, which introduces a robust Hardware-Aware Proxy model to enhance prediction accuracy and hardware versatility.
Abstract
Quantized Neural Networks~(QNN) with low-bitwidth data have proven promising in efficient storage and computation on edge devices. To mitigate accuracy degradation while maximizing speedup, layer-wise mixed-precision quantization~(MPQ) becomes a popular solution. However, existing algorithms for exploring MPQ schemes are limited in flexibility and efficiency. Comprehending the complex impacts of different MPQ schemes on post-training quantization and quantization-aware training results is a challenge for conventional methods. Furthermore, an end-to-end framework for the optimization and deployment of MPQ models is missing in existing work. To address these challenges, we propose the MiCo framework, a holistic MPQ exploration and deployment framework for edge AI applications. The framework adopts a novel optimization algorithm to search for accuracy-optimal quantization configurations under strict latency constraints. We further extended the framework to MiCoPro, which introduces a robust Hardware-Aware Proxy (HAP) model to enhance prediction accuracy and hardware versatility. By leveraging target-specific latency modeling, MiCoPro enables rapid exploration and direct deployment from PyTorch models to bare-metal C code. We demonstrate the versatility of our framework on both the BitFusion accelerator and SIMD-extended RISC-V processors, achieving up to 40\% of latency reduction with less than 3\% of accuracy drop.
Conventional compute-in-memory (CIM) deployment flows usually assume a fixed crossbar geometry, although DNN layers often have different channel counts and matrix shapes. The resulting shape mismatch leaves part of the array capacity unused and can increase the number of split-and-transfer operations during inference....
Kaiwen Deng, Sifan Sun, Hanjie Liu et al.· IEEE Non-Volatile Memory Sys...· 0 citations
This paper proposes a streamable, client-server NVC architecture featuring a novel Mixed Precision (FP16/FP32) strategy, and demonstrates that this approach successfully eliminates intra-generation fragmentation and substantially broadens cross-die interoperability, achieving seamless cross-generation decodability for...
Kasidis Arunruangsirilert, He-Ming Sun, J. Katto· 0 citations
Analog compute-in-memory (CIM) has recently emerged as a novel paradigm for artificial intelligence compute, but the efficiency of CIM is heavily bottlenecked by the energy and area overhead of analog-to-digital (ADC). While replacing high-precision ADCs with 1-bit conversion significantly reduces peripheral overhead,...
Wei-Wei Zhao, Sohan Salahuddin Mugdho, Cheng Wang et al.· Midwest Symposium on Circuit...· 0 citations
Faster Flash Decoding (FFD) is presented, a novel hardware-algorithm co-design framework designed to break the memory wall in long-context decoding and introduces the top-delta strategy, which dynamically filters blocks to achieve distribution-adaptive sparsity without global synchronization.
Zhigeng Liu, Zhiyuan Ning, Rui-Xiao Li et al.· 4 citations
OptiPrime is introduced, a protocol-hardware co-optimization framework for efficient private DNN inference that features a novel HE protocol for convolutions that substantially reduces the number of transmitted output ciphertexts and mitigates the network communication bottleneck.
PUMA combines offline layer/layer-chain PMU characterization with online GPU PMU observations to identify execution characteristics associated with compute-bound, memory-bound, and bursty phases and achieves a lower energy-delay product than the existing governor across all evaluated workloads.
W. Chang, Seung-Ryeol Ohk, Young-Jin Kim· IEEE Access· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.