KBBQ is introduced, which parameterizes the extent to which a transform approaches this ceiling, and outperforms the prior state of the art without additional deployment-time computation.
Abstract
We develop a second-order theory of quantization noise in matrix multiplication in which the quantization format is characterized by the variance it assigns to each element. The constant variance profile of integer quantization recovers existing integer-noise theory, while the multiplicative profile of floating-point rounding reduces the data dependence to a scalar, the participation factor $\kappa$, yielding a closed-form signal-to-noise-ratio law. The resulting functional also admits a closed-form upper bound $\kappa^{*}$ that no function-preserving linear transform can exceed and that is attained by a recent state-of-the-art method. Building on this analysis, we introduce KBBQ (\textbf{K}appa-\textbf{B}raked \textbf{B}lockwise \textbf{Q}uantization), which parameterizes the extent to which a transform approaches this ceiling. At W4A4, across four base models and two FP4 formats, KBBQ outperforms the prior state of the art without additional deployment-time computation.
We consider the problem of minimizing error in quantized matrix multiplication $C=AB$. Scalar quantization of the factors introduces rounding errors whose scale depends on the maximum absolute entries -- the ranges -- of their rows and columns. These ranges determine the quantization grid steps. To reduce the error, we...
Piyush Sao, N. Miniskar, Pedro Valero-Lara et al.· 1 citation
Quantization-Aware Pre-Training (QAPT) can increase the inference efficiency of DNNs, but a problematic behaviour known as rounding boundary weight oscillation can introduce detrimental noise into the training process and significantly reduce convergence speed. While existing methods can reduce this detrimental noise,...
Post-training quantization (PTQ) methods in the GPTQ family minimize a layer-wise reconstruction error on a uniform grid whose scale must be chosen; the common max-based choice degrades sharply at low bit-widths. We study how sensitive this objective is to the scale. For a layer with i.i.d. Gaussian weights and calibra...
Jonas von Berg, Massimiliano Datres, Carlo Kneissl et al.· 0 citations
This work identifies and formalizes the principle that organizes the Great Inversion, the Great Inversion: allocation-flexible coding rewards energy concentration, whereas the grouped shared-scale quantization a deployed matrix instruction performs rewards within-group flattening.
This work proposes Subsampled Stochastic TurboQuant (SSTQ), a framework that combines overcomplete equal-norm tight frames, coordinate subsampling, and privacy-aware one-dimensional quantization and empirically evaluates SSTQ against established baselines on federated learning tasks using CIFAR-10 and Fashion-MNIST, de...
Adel Javanmard, David P. Woodruff, V. Mirrokni· 0 citations
We revisit Kashin-decomposition-based weight quantization for large language models and propose an improved algorithm with stronger convergence properties and structured, efficient orthogonal transforms. The method retains the core factorization of each weight into two components -- one with bounded infinity norm and t...
Daria Cherniuk, A. Rudikov, B. Kashin et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 29, 2026
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.