Preprint
Jul 2026
PolyQ: Codesigning End-to-End Quantization Framework for Scalable Edge CPU LLM Inference
PolyQ, a CPU-oriented compiler/quantization co-design for activation-aware channel-wise bit allocation under a user-specified average-bit budget, shows that fractional-bit CPU deployment is practical, predictable, and energy-efficient across diverse edge targets.
Hyunwoo Oh, Suyeon Jang, Hanning Chen et al.
· 0 citations