Skip to content
Preprint

Revisiting TD Target Aggregation under Uncertainty in Q-Learning

Aug 2026 · 0 citations · 60 references
Computer Science

TL;DR

The proposed SADQ is a simple modification to Q-learning that regularizes how the TD target is formed, and consistently improves training stability across classical control tasks, real-world vector-based environments, and Atari benchmarks when compared to strong DQN variants.

Abstract

Deep Q-Networks (DQNs) learn value functions through bootstrapped temporal-difference updates, where future returns are approximated using a greedy maximization over next-state action values. While effective, this aggregation rule is inherently sensitive to estimation noise: when Q-values are uncertain, the maximization operator deterministically favors the largest estimate, regardless of its reliability, leading to amplified errors through bootstrapping. In this work, we propose the \textbf{S}uccessor Rollout \textbf{A}ggregation \textbf{D}eep \textbf{Q}-Network (SADQ), a simple modification to Q-learning that regularizes how the TD target is formed. SADQ uses one-step rollout predictions from a learned dynamics model to guide the comparison among candidate next-state actions, introducing additional structure into the aggregation step without altering the underlying learning framework. The resulting mixed Bellman update attenuates unreliable maxima while preserving the standard fixed point under diminishing model error. We provide theoretical analysis showing that SADQ reduces bootstrap-induced overestimation in a pointwise manner. Empirically, SADQ consistently improves training stability across classical control tasks, real-world vector-based environments, and Atari benchmarks when compared to strong DQN variants.

View source

Similar papers

Preprint Aug 2026

Adaptive Finite-Budget Training for CVaR Risk-Aware Q-Learning

Results demonstrate that adaptive finite-budget training design, applied solely to the training procedure without altering the risk objective, can materially improve the reliability and risk-adjusted performance of risk-aware Q-learning in financial applications.

Yifan Wu, Junjie Lei, Wenjie Huang · 0 citations
Jul 2026

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning

Deep off-policy reinforcement learning algorithms for continuous control typically rely on neural value function approximation to guide policy improvement. However, temporal-difference (TD) learning introduces noisy targets, resulting in non-stationary optimization, while greedy policy updates amplify early-stage estim...

Gong Gao, Xiao Lai, Zi-Qi Xie et al. · 0 citations
#artificial intelligence Preprint Sep 2026

On BatchNorm Forward Modes in Value-Based Reinforcement Learning

Batch normalization (BN) substantially improves sample efficiency in continuous-control actor-critic methods such as CrossQ, yet recent studies report performance degradation in discrete-action value learning on Atari. These failures are surprising because discrete Q-networks lack the action-input distribution mismatch...

Daniel Palenicek, Mikael Henaff, Scott Fujimoto et al. · 0 citations
Preprint Aug 2026

Start Classifying: Categorical Critics for LLM Reinforcement Learning

Proximal Policy Optimization (PPO) for large language models typically trains its critic by mean-squared-error (MSE) regression on scalar value targets. Although scalar MSE is statistically valid for estimating the conditional expected return, sparse binary rewards in reinforcement learning with verifiable rewards (RLV...

Zhi-Jian Zhou, Long Li, Xuan Zhang et al. · 2 citations · ⚡2
Preprint Aug 2026

Offline Deep Q* Estimation with Diffusion Models

In offline RL, estimating the optimal action-value function $Q^*$ can be formulated as solving the optimal Bellman equation based solely on offline observations. A fundamental challenge is that the reward function and transition kernel are unknown, so the optimal Bellman operator is not directly observable from data. T...

Xiao-Hong Chen, Yu-Ling Jiao, Lican Kang et al. · 0 citations
#machine learning Preprint Sep 2026

A Convergence Framework for Deep $V$-Learning: Error Propagation and Sharp Action-Gap Bounds

We establish convergence bounds for deep $V$-learning with horizon $H$. The algorithm fits a scalar value function to targets from executed transitions and selects actions using a predictive model and the value function. For current observed-successor targets with fresh true-kernel outcomes, the conditional mean is $\m...

Yury Kolomeytsev · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.