Mean-Field PhiBE: Continuous-Time Mean-Field Reinforcement Learning from Discrete-Time Data
A first-order consistency estimate is proved showing that the value induced by an optimal MF-PhiBE policy approximates the optimal continuous-time value as the observation time step vanishes, and a policy-gradient theorem for entropy-regularized randomized feedback policies is derived.