Skip to content

Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents

Jul 2026 · arXiv.org · Vol abs/2607.27275 · 1 citation · 16 references
Computer Science

TL;DR

This claim for multi-turn, tool-calling agents, where it now matters most, is tested for post-training quantization to 4-bit weights and diagnostics, the per-channel error rate and success under a shrinking budget come from logs benchmarks already collect.

Abstract

Post-training quantization to 4-bit weights is widely reported to be nearly lossless. We test this claim for multi-turn, tool-calling agents, where it now matters most. On $\tau^2$-bench, across two open-weight model families in dense and MoE variants and two domains (eight cells, 456 episodes each, at 16-, 8-, and 4-bit weights), quantization indeed looks free on the standard metric. No cell shows a score change that survives multiple-comparison correction, and in the cell that carries the largest process damage, equivalence testing bounds the change within $\pm$7.5 points. The process tells a different story. Quantization amplifies the failure the model already exhibits at full precision (tool-name hallucination in telecom, with the same directional trend in retail entity errors) by up to 2.5$\times$ in volume (+17.6 points per task), while creating essentially no new failures. The failure set is the same at every precision (rank correlation $\geq$ 0.94, 0.18% novel events). The score stays flat because the benchmark's ten-error budget absorbs the extra failures. Shrinking the budget to two errors re-exposes a score gap of 17 points, and it does so only in the one cell where quantization added error volume, exactly as the masking account predicts. A targeted error-repair prompt, run for five telecom models at every precision, removes the damage exactly and only where it lives. Both diagnostics, the per-channel error rate and success under a shrinking budget, come from logs benchmarks already collect; we suggest reporting them alongside task reward.

View source

Similar papers

#machine learning Preprint Sep 2026

The Structure of Quantization Damage in LLMs: Why the Next Bit Should Be Spent Globally

Post-training quantization (PTQ) is widely used to reduce the cost of serving large language models (LLMs), but its accuracy cost is uneven and is often tuned per model. We study where quantization damage occurs and how to allocate a small additional precision budget. Using causal mixed-precision intervention as ground...

Jun-Hao Hu, S. Ramachandran · 1 citation
#natural language process... Preprint Aug 2026

Which Decisions Low-Bit Quantization Breaks, and How to Predict Them

This work tracks quantization across 16 models from 8 families under round-to-nearest, seven under AWQ, two under GPTQ and one under GGUF, at 8 down to 2 bits, and measures the margin, the picked option's score minus its best alternative's, which removes the protection a large margin affords.

Zekun Wu, Swati Dhiman, A. Koshiyama · 1 citation
#artificial intelligence Preprint Sep 2026

Quantization Amplifies Determinism, Not Bias: Scale-Dependent Behavioral Effects of Serving-Time Weight Compression

This work serves three checkpoints at three weight precisions, holding the hardware, software, and sampling configuration constant, and collects approximately 71,000 completions paired by prompt and seed across two custom, leak-checked prompt batteries.

Dachi Kurtskhalia · 0 citations
Jul 2026

Bigger or Cheaper? Scale and Quantization Effects on Uncertainty Signals in Vision-Language Models Under Image Degradation

This work measures, on identical inputs, how model scale and 4-bit quantization affect two confidence signals in the Qwen2-VL family: the confidence a model states in natural language, and its own mean token probability over the answer it generates, and argues that error-detection AUROC is the metric that exposes the d...

M. M. Asif Ferdous · 0 citations
Jul 2026

QuantiBias: Benchmarking Quantization-Induced Bias in LLMs

A benchmark that pairs a generative, multilingual stereotype probe with the refusal and multiple-choice controls that isolate open-ended generation, contrasts each build with and without reasoning, and rates the content severity of what it generates.

Emilio Ferrara · 0 citations
Preprint Aug 2026

SchurQuant: Groupwise Discrete Optimization for Layer-Wise LLM Quantization

SCHUROPT is introduced, which analytically eliminates the suffix's optimal continuous response, yielding an exact groupwise quadratic with Schur-complement curvature, and achieves the highest mean zero-shot accuracy among the evaluated backpropagation free PTQ baselines.

Gunjun Lee, Sehwan Son, Younjoo Lee et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.