Preprint
Aug 2026
Revisiting TD Target Aggregation under Uncertainty in Q-Learning
The proposed SADQ is a simple modification to Q-learning that regularizes how the TD target is formed, and consistently improves training stability across classical control tasks, real-world vector-based environments, and Atari benchmarks when compared to strong DQN variants.
Li-Peng Zu, Xiaonan Zhang
· 0 citations