Skip to content
Preprint

Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration

Aug 2026 · 0 citations · 79 references
Computer Science

TL;DR

This work proposes maximizing the entropy of predictive distributions as the calibration objective, which directly targets overconfidence by discouraging overly concentrated predictions, and adopts an efficient first-order approximation that avoids explicit second-order computation.

Abstract

Preference alignment often makes large language models (LLMs) overconfident and poorly calibrated. Traditional post-hoc temperature scaling is inherently domain-dependent: a temperature fitted on one domain does not generalize across domains. This motivates us to modify model parameters during training to improve calibration. We propose maximizing the entropy of predictive distributions as the calibration objective, which directly targets overconfidence by discouraging overly concentrated predictions. Inspired by temperature scaling, we realize this through a bilevel optimization formulation, where the lower level trains the model under a parametric loss and the upper level selects loss hyperparameters to maximize entropy. To make the framework practical at LLM scale, we adopt an efficient first-order approximation that avoids explicit second-order computation. Across both multiple-choice and open-ended generative question answering, experiments demonstrate that our method yields well-calibrated LLMs with particular advantages in out-of-domain generalization.

View source

Similar papers

Preprint Aug 2026

Weightless Fine-Tuning: Personalizing LLMs via Logit-Space Transport

Weightless Fine-Tuning is proposed, a training-free decoding-time method that approximates the distributional effect of SFT without weight updates and achieves the best average performance across datasets, matches or exceeds SFT on individual tasks, and outperforms other lightweight baselines on average.

Bo-Han Zhang, An-Qi Ni, Yi-Xin Wang et al. · 0 citations
#machine learning Preprint Sep 2026

Bridging the Confidence Gap: Temperature Scaling for Calibrating Test-Time Prompt Tuning

Test-time prompt tuning (TPT) enables adaptation on a single test instance, achieving improved accuracy but often sacrificing calibration performance. Most existing calibration methods introduce additional regularization terms to promote dispersion across text embeddings and reduce calibration error, yet these methods...

Yu-Wei Liang, Jian Liang, Dapeng Hu et al. · 0 citations
Preprint Aug 2026

Target-Aware Calibration Data Selection for Preserving Uncertainty in Quantized Language Models

DPQ is introduced, a lightweight pre-quantization recipe family that uses full-precision predictions to construct target-aligned calibration mixtures of high-doubt examples and generic anchors that better preserve broad multiple-choice QA behavior.

Zhen Yang, Sizai Hou, Kai-Wen Zheng et al. · 0 citations
#machine learning Preprint Sep 2026

LILA: Calibration-Free Structured Pruning of Large Language Models via Latent Spectral Geometry

Structured pruning of large language models (LLMs) offers hardware-efficient compression, yet existing methods require calibration data, gradient computation, or large auxiliary policy networks at pruning time. LILA (\emph{Latent-Informed Layer Analysis}) scores neuron importance via the Kolmogorov--Smirnov (KS) distan...

Sankar Behera, D. Singh, Anshika Agnihotri et al. · 0 citations
#natural language process... Preprint Sep 2026

From Rollouts to Recipes: Self-Contained Post-Training for LLMs

Self-Routing is proposed, a behavior-conditioned post-training framework that uses rollout correctness and confidence to decide how each sample should be optimized, allowing training to adapt without external teachers, extra annotations, or additional sampling.

Yifei Li, Lingling Zhang, Mu-Ye Huang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Distribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts

Mixture-of-Experts (MoE) architectures have emerged as a powerful paradigm for scaling model capacity while preserving efficient inference in large foundation models. However, most MoE models use a fixed top-$k$ expert selection policy, assigning the same expert budget to every token even when fewer experts may be suff...

Dohyeon Kim, Bedionita Soro, Sung Ju Hwang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.