Skip to content
Preprint

Bellman Calibration for Marginalized Importance Weighting in Offline Reinforcement Learning

Aug 2026 · 0 citations · 87 references
Computer Science Mathematics

TL;DR

The introduction of isotonic Bellman calibration, a one-dimensional, model-agnostic post-processing method that reduces residual occupancy-balance violations while preserving the ranking information in any initial occupancy-ratio estimate, and establishes finite-sample calibration guarantees and a KL oracle inequality relative to the best monotone transformation of the initial estimate.

Abstract

Marginalized importance weighting evaluates a target policy by reweighting offline state-action samples with its discounted occupancy ratio, characterized by an adjoint Bellman equation. Existing minimax, primal-dual, and fitted fixed-point estimators can leave residual occupancy-balance violations because of function-class approximation, regularization, or incomplete optimization. These violations are difficult to diagnose and reduce because the objectives generally lack a direct supervised validation loss for hyperparameter tuning, model selection, and early stopping. We introduce isotonic Bellman calibration, a one-dimensional, model-agnostic post-processing method that reduces these violations while preserving the ranking information in any initial occupancy-ratio estimate. The method corrects the estimate's scale and shape by applying fitted occupancy-ratio evaluation (FORE) over a one-dimensional class of nondecreasing transformations. We characterize Bellman calibration as a conditional fixed-point property equivalent to occupancy-balance against every test function of the calibrated ratio. More generally, we derive a calibration-refinement bound showing that any fitted ratio with small calibration error performs nearly as well as the best post-processing based on its fitted values. For isotonic Bellman calibration, we establish finite-sample calibration guarantees and a KL oracle inequality relative to the best monotone transformation of the initial estimate. Consequently, isotonic Bellman calibration achieves small calibration error and KL risk within statistical error of the best monotone correction, with guarantees for downstream target-occupancy functionals, including policy-value estimation.

View source

Similar papers

Preprint Aug 2026

Duality and Error for Predictively Oriented Inference

This work derives a finite-dimensional dual formulation of PrO inference that separates sampling fluctuation, approximation under a divergence budget, regularization, and numerical optimization error and uses an exactly solvable categorical example to show that predictive-risk convergence can imply convergence to a uni...

Aurya Javeed, D. Kouri, Teresa Portone et al. · 0 citations
Preprint Aug 2026

ReBRAC-v2: The Return of the King

ReBRAC-v2 is introduced, which directly trains an exact-likelihood normalizing flow as the RL actor, combines likelihood, MSE, and MAE behavior regularization, and integrates a classification-based residual critic, staged optimization, and multi-sample test-time action selection, and ranks first in eight categories.

Denis Tarasov, Robert K. Katzschmann · 0 citations
Preprint Aug 2026

Start Classifying: Categorical Critics for LLM Reinforcement Learning

Proximal Policy Optimization (PPO) for large language models typically trains its critic by mean-squared-error (MSE) regression on scalar value targets. Although scalar MSE is statistically valid for estimating the conditional expected return, sparse binary rewards in reinforcement learning with verifiable rewards (RLV...

Zhi-Jian Zhou, Long Li, Xuan Zhang et al. · 2 citations · ⚡2
Jul 2026

ISO: An RLVR-Native Optimization Stack

Reinforcement learning with verifiable rewards (RLVR) is rapidly advancing the reasoning capabilities of language models, yet the optimization layer that converts reward feedback into weight-space updates remains poorly understood. Building on our prior analysis (Zhu et al., 2025), we study this missing layer through t...

Hanqing Zhu, Wenyan Cong, Zhizhou Sha et al. · 0 citations
Preprint Aug 2026

Adaptive Finite-Budget Training for CVaR Risk-Aware Q-Learning

Results demonstrate that adaptive finite-budget training design, applied solely to the training procedure without altering the risk objective, can materially improve the reliability and risk-adjusted performance of risk-aware Q-learning in financial applications.

Yifan Wu, Junjie Lei, Wenjie Huang · 0 citations
#machine learning Preprint Sep 2026

Online Self-Weighted Fine-Tuning

Online Self-Weighted Fine-Tuning is proposed, a simple method that augments SFT with online, trajectory-level weighting and offers a favorable compute-performance trade-off as a practical approach for fine-tuning small-to-medium LLMs on binary-verifiable reasoning tasks with only 2 online rollouts.

Hai-Quan Wen, Yiwei He, Bei Peng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.