SSQT: A Hardware-Friendly Fusion Compression Framework of Structured Sparsification and Sensitivity-Driven Quantization for Large-Scale Language Models
The results show that compression and execution layout must be co-designed to convert model-size reduction into practical parallel inference acceleration, and that compression and execution layout must be co-designed to convert model-size reduction into practical parallel inference acceleration.