Preprint
Jul 2026
Learning the Supports for Categorical Critic in Reinforcement Learning
This work investigates the Gaussian Histogram Loss (HL-Gauss), a recent approach that reframes value estimation as classification by encoding each scalar Bellman target as a Gaussian-smoothed categorical target, and derives an objective that forms an upper bound on the mean-squared Bellman error.
Jen-Yen Chang, Takayuki Osa, Tatsuya Harada
· 0 citations