G$^2$PTQ is presented, a unified PTQ framework with Generalized Gradient Compensation that integrates both first- and second-order information under a globally supervised, block-wise optimization objective and enables better alignment with the full-precision model, outperforming state-of-the-art baselines.
Abstract
Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large language models (LLMs) without retraining. GPTQ-based methods have become the de facto standard, yet they suffer from two complementary limitations. Methods with local, layer-wise objectives lack global supervision; while methods with global objectives fix their Hessian estimates at the start and ignore first-order gradients, so their guidance grows stale as quantization proceeds. This paper presents G$^2$PTQ, a unified PTQ framework with Generalized Gradient Compensation that integrates both first- and second-order information under a globally supervised, block-wise optimization objective. By refreshing gradient and Hessian estimates before quantizing each Transformer block, G$^2$PTQ avoids the staleness of prior global methods. Furthermore, to stabilize the exact first-order compensation, we introduce a trust-region scaling mechanism that dynamically bounds the gradient step to prevent exploding weight updates. Finally, we derive efficient implementations for block-wise Hessian approximation and exact gradient compensation. Experimental results on various model families and bit-widths demonstrate that G$^2$PTQ enables better alignment with the full-precision model, outperforming state-of-the-art baselines. Code is available at: https://github.com/G2PTQ/G2PTQ.
Real-Q (Real-time E2E-loss Aligned LLM Quantization), a novel PTQ paradigm that breaks this compromise: instead of diluting the objective for the sake of analytic tractability, REAL-Q targets an end-to-end-aligned surrogate of the global loss and refines it via fine-grained, dynamic Block-wise Gradient Descent applied...
Qian Zhang, Yao-Ming Li, Zheng Tan et al.· 0 citations
OJBKQ is proposed, a layer-wise PTQ method that formulates weight quantization as a joint optimization problem over activations and weights, yielding a multiple-right-hand-side box-constrained integer least squares (BILS) problem per layer.
Xin-Yu Wang, Zi-Yu Zhao, Peng Lu et al.· 0 citations
SCHUROPT is introduced, which analytically eliminates the suffix's optimal continuous response, yielding an exact groupwise quadratic with Schur-complement curvature, and achieves the highest mean zero-shot accuracy among the evaluated backpropagation free PTQ baselines.
Gunjun Lee, Sehwan Son, Younjoo Lee et al.· 0 citations
Quantization-Aware Pre-Training (QAPT) can increase the inference efficiency of DNNs, but a problematic behaviour known as rounding boundary weight oscillation can introduce detrimental noise into the training process and significantly reduce convergence speed. While existing methods can reduce this detrimental noise,...
Post-training quantization (PTQ) efficiently compresses Vision Transformers (ViTs) without retraining, yet suffers severe accuracy degradation at low bit-widths. Existing optimization-based PTQ methods guide block reconstruction via either soft logits or second-order Hessian proxies. Logit supervision is prone to overf...
Pei Sun, Guang-Qi Liang, Jin-Nian Tong et al.· 0 citations
This paper introduces Gradient Enhancement Task Aware Post-training Quantization, i.e., GTAQ, to address the generalization issue of Large Language Models, and extensively evaluates the LLaMA family of language models on WikiText, C4, and MMLU.
Yi-Hua Shao, Yang-Yang Gu, Min-Xi Yan et al.· Proceedings of the Thirty-Fi...· 1 citation