Adaptive-optimal control of vehicle lateral dynamics via integral-reinforcement-learning actor–critic policy iteration
Simulation results show that the IRL actor–critic controller matches or improves upon the fixed-gain baseline after a plant change while, unlike the certainty-equivalence scheme, requiring no continuous probing/dither signal for identifiability—avoiding the associated persistent tracking-error penalty.