Bellman Calibration for Marginalized Importance Weighting in Offline Reinforcement Learning
The introduction of isotonic Bellman calibration, a one-dimensional, model-agnostic post-processing method that reduces residual occupancy-balance violations while preserving the ranking information in any initial occupancy-ratio estimate, and establishes finite-sample calibration guarantees and a KL oracle inequality...