#machine learning
Jun 2026
Mean-Field PhiBE: Continuous-Time Mean-Field Reinforcement Learning from Discrete-Time Data
A first-order consistency estimate is proved showing that the value induced by an optimal MF-PhiBE policy approximates the optimal continuous-time value as the observation time step vanishes, and a policy-gradient theorem for entropy-regularized randomized feedback policies is derived.
Erhan Bayraktar, Mart'in Hern'andez, Qin-Xin Yan et al.
· arXiv.org · 1 citation