SSQT: A Hardware-Friendly Fusion Compression Framework of Structured Sparsification and Sensitivity-Driven Quantization for Large-Scale Language Models
Large language models (LLMs) are often memory-bandwidth bound during autoregressive decoding, so reducing weight storage does not automatically produce parallel speedup when the compressed representation is irregular. We present SSQT, a post-training framework that jointly applies hardware-aligned structured sparsification, sensitivity-driven outlier preservation, and sparse-aligned low-bit group quantization. SSQT estimates parameter importance from calibration data with a diagonal empirical-Fisher approximation, avoiding construction of the full Hessian, and, in the default 4-bit configuration, stores fewer than 1% high-sensitivity weights on a separate FP16 residual path. The remaining weights are packed in regular attention blocks and contiguous feed-forward-network channel groups; quantization metadata is secondarily quantized and decoded inside the matrix-multiplication tile rather than by globally expanding the model to FP16. Experiments on Llama 2 and Falcon models report task quality, calibration cost, packed storage, latency regularity, cross-GPU results, and hardware counters. In the default 4-bit configuration, SSQT uses 24.3% of the FP16 model-memory footprint on Llama 2-13B, keeps the relative WikiText2 perplexity increase at 4.8%, and improves the per-sequence decoding rate by up to 2.33 × on the A100 tensor-parallel configuration. On Llama 2-13B, Tensor Core utilization rises from 28.7% to 62.4%, memory-bandwidth utilization falls from 92.3% to 41.8%, and pipeline stalls fall from 34.6% to 11.0%. These results show that compression and execution layout must be co-designed to convert model-size reduction into practical parallel inference acceleration.