Skip to content

Greedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence in LLM Inference

Sep 2026 · 0 citations · 43 references
Computer Science

TL;DR

An empirical error-propagation analysis is developed and finds that 22 layers of accumulated body error do not distinguish flipping from non-flipping steps; the outcome depends primarily on the top-two logit margin at the LM head relative to the directional perturbation between the top-two candidates.

Abstract

Greedy decoding from large language models is commonly treated as deterministic. We show it is not precision-invariant: the same model, prompt, and decoding algorithm produce different outputs in BF16 versus FP16 on identical hardware. Across our evaluations of six models (1.1B-7B parameters, four families; divergence additionally characterised at 12B) and three benchmarks, 49-100\% of prompts diverge; a single token flip often cascades into trajectory-level divergence. We develop an empirical error-propagation analysis and find that 22 layers of accumulated body error do not distinguish flipping from non-flipping steps; the outcome depends primarily on the top-two logit margin at the LM head relative to the directional perturbation between the top-two candidates. The analysis makes five testable predictions about intervention outcomes, including that applying more FP32 compute (broader scope) makes agreement worse. The experiments match all five predictions. The best-performing low-overhead intervention we evaluate, selective FP32 LM head recomputation, triggered only when the margin falls below a threshold, delivers +22-36 pp exact agreement on A10G (+12-21 pp on L4 and A100) at less than 4\% latency overhead in low-batch (batch size<=4) single-stream inference. We map the applicability boundary across six models and four batch sizes, and hypothesise that training-time precision stability is a determining factor. The method is a partial mitigation rather than a universal determinism guarantee: its benefit vanishes when body-originated error dominates, including at batch size>=8 and under end-to-end FP8 in our tests.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Evaluating Losslessness in Speculative Decoding Under Finite-Precision Inference

Lossless speculative decoding is typically defined at the algorithmic level: a speculative procedure proposes multiple tokens and a verification procedure is designed to preserve the output trajectory of an autoregressive reference model exactly. In practical neural inference, however, this guarantee is implemented usi...

Ilya Koziev, Leonid S. Sinev, I. Oseledets · 0 citations
#artificial intelligence Preprint Sep 2026

Accuracy is Not Enough: A Divergence-Based Approach to Evaluate Fidelity Loss in Quantized LLMs

Deployment of Large Language Models (LLMs) on memory-constrained edge devices relies heavily on aggressive post-training quantization. However, evaluating these models is largely based on zero-shot task accuracy, which depends solely on argmax predictions and is insensitive to changes in the underlying predictive distr...

Shahzeb Qamar, L. Sparrenberg, Christian Bauckhage et al. · 1 citation
#machine learning Preprint Sep 2026

The Structure of Quantization Damage in LLMs: Why the Next Bit Should Be Spent Globally

Post-training quantization (PTQ) is widely used to reduce the cost of serving large language models (LLMs), but its accuracy cost is uneven and is often tuned per model. We study where quantization damage occurs and how to allocate a small additional precision budget. Using causal mixed-precision intervention as ground...

Jun-Hao Hu, S. Ramachandran · 1 citation
Preprint Sep 2026

Contextual Utility of Quantization Moves in Extreme Low-Bit LLMs

Post-training quantizers select finite code changes using reconstruction proxies or local loss approximations, but the utility of a quantization move depends on the state through which it is executed. We identify two sources of this contextual dependence. First, the displacement of the move matters: evaluating the grad...

Wen-Xuan Xiao, Xu Cao · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond Truncation: Rethinking LLM Decoding as Ensemble Pruning

We introduce Mahalanobis-Ensemble Decoding (ME-Decoding), a novel Large Language Model (LLM) decoding framework that frames candidate token selection as ensemble pruning. Existing selection strategies rely predominantly on scalar probabilities, ignoring geometric semantic relationships and causing candidate redundancy....

Dun-Yao Xue, Cheng-Shuo Du, Zheng-Bo Wang et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.