Skip to content

KBBQ: A Predictive Noise Law and the Limits of Spectrum Flattening in FP4 Quantization

Sep 2026 · 0 citations · 34 references
Computer Science

TL;DR

KBBQ is introduced, which parameterizes the extent to which a transform approaches this ceiling, and outperforms the prior state of the art without additional deployment-time computation.

Abstract

We develop a second-order theory of quantization noise in matrix multiplication in which the quantization format is characterized by the variance it assigns to each element. The constant variance profile of integer quantization recovers existing integer-noise theory, while the multiplicative profile of floating-point rounding reduces the data dependence to a scalar, the participation factor $\kappa$, yielding a closed-form signal-to-noise-ratio law. The resulting functional also admits a closed-form upper bound $\kappa^{*}$ that no function-preserving linear transform can exceed and that is attained by a recent state-of-the-art method. Building on this analysis, we introduce KBBQ (\textbf{K}appa-\textbf{B}raked \textbf{B}lockwise \textbf{Q}uantization), which parameterizes the extent to which a transform approaches this ceiling. At W4A4, across four base models and two FP4 formats, KBBQ outperforms the prior state of the art without additional deployment-time computation.

View source

Similar papers

#machine learning Preprint Sep 2026

A Nuclear-Norm Lower Bound for Dithered Scalar Quantization of Matrix Products

We consider the problem of minimizing error in quantized matrix multiplication $C=AB$. Scalar quantization of the factors introduces rounding errors whose scale depends on the maximum absolute entries -- the ranges -- of their rows and columns. These ranges determine the quantization grid steps. To reduce the error, we...

Piyush Sao, N. Miniskar, Pedro Valero-Lara et al. · 1 citation
#machine learning Preprint Sep 2026

Quantization-Aware Pre-Training with Constrained Empirical Weight Distribution

Quantization-Aware Pre-Training (QAPT) can increase the inference efficiency of DNNs, but a problematic behaviour known as rounding boundary weight oscillation can introduce detrimental noise into the training process and significantly reduce convergence speed. While existing methods can reduce this detrimental noise,...

Ning-Feng Yang, T. Aamodt · 0 citations
#machine learning Preprint Sep 2026

Scale Sensitivity in Low-Bit Post-Training Quantization: Curvature of the Quantization Error Landscape

Post-training quantization (PTQ) methods in the GPTQ family minimize a layer-wise reconstruction error on a uniform grid whose scale must be chosen; the common max-based choice degrades sharply at low bit-widths. We study how sensitive this objective is to the scale. For a layer with i.i.d. Gaussian weights and calibra...

Jonas von Berg, Massimiliano Datres, Carlo Kneissl et al. · 0 citations
Review Aug 2026

Transforms for LLM Quantization: The Great Inversion and Format Co-Design

This work identifies and formalizes the principle that organizes the Great Inversion, the Great Inversion: allocation-flexible coding rewards energy concentration, whereas the grouped shared-scale quantization a deployed matrix instruction performs rewards within-group flattening.

Ehsan Jokar · 0 citations
#machine learning Preprint Aug 2026

SSTQ:Privacy-Preserving Vector Quantization via Subsampled Stochastic TurboQuant

This work proposes Subsampled Stochastic TurboQuant (SSTQ), a framework that combines overcomplete equal-norm tight frames, coordinate subsampling, and privacy-aware one-dimensional quantization and empirically evaluates SSTQ against established baselines on federated learning tasks using CIFAR-10 and Fashion-MNIST, de...

Adel Javanmard, David P. Woodruff, V. Mirrokni · 0 citations
#natural language process... Preprint Sep 2026

Structured Transforms for Low-Overhead Quantization of Language Models

We revisit Kashin-decomposition-based weight quantization for large language models and propose an improved algorithm with stronger convergence properties and structured, efficient orthogonal transforms. The method retains the core factorization of each weight into two components -- one with bounded infinity norm and t...

Daria Cherniuk, A. Rudikov, B. Kashin et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.