Preprint
Sep 2026
Vortex: Bridging Extreme Compression and Efficient LLM Inference
This study addresses challenges with Vortex, an architecture compatible with systolic-array-based accelerators with minimal hardware overhead, bridging the gap between extreme compression and efficient inference, and proposes codebook-wise contextual sparsity to align with VQ execution.
Haoxuan Shan, Cong Guo, Bo-Wen Duan et al.
· 0 citations