Recent advances in neural network design are integrated: observation and feature normalization, weight normalization, and modeling of distributional returns with an entropy-regularized MORL algorithm, demonstrating that these changes substantially improve the quality of the produced solution sets without requiring major changes to the underlying algorithm.
Abstract
Recent advances in deep reinforcement learning (RL) have shown that improving neural network architectures can yield substantial gains in sample efficiency and asymptotic performance without altering the underlying algorithms. In contrast, work on multi-objective reinforcement learning (MORL), which aims to discover a set of policies that balance trade-offs among conflicting objectives, has predominantly focused on algorithmic innovations, leaving the area of architectures underexplored. While the optimal policies and value functions can differ significantly depending on the trade-offs, MORL algorithms commonly represent them with simple feedforward networks conditioned on the trade-off. This raises the question of whether the performance of the algorithms could be improved with more expressive function approximators. In this paper, we integrate recent advances in neural network design: (i) observation and feature normalization, (ii) weight normalization, and (iii) modeling of distributional returns with an entropy-regularized MORL algorithm. The empirical results across standard continuous control benchmarks demonstrate that these changes substantially improve the quality of the produced solution sets without requiring major changes to the underlying algorithm.
Deep reinforcement learning (DRL) relies on neural networks trained under temporally correlated experience, evolving replay distributions, and bootstrapped targets, conditions that can destabilize neural function approximation. This study investigates controllable orthogonalization as a network-level mechanism for impr...
Jia Yu, Dun Su, Hua Cui et al.· Informatics· 0 citations
Reinforcement learning (RL) has become a cornerstone for improving the reasoning capabilities of large language models (LLMs), but the need for on-policy data substantially limits training efficiency. Reusing off-policy data through importance sampling (IS) can improve efficiency but introduce considerable instability....
Batch normalization (BN) substantially improves sample efficiency in continuous-control actor-critic methods such as CrossQ, yet recent studies report performance degradation in discrete-action value learning on Atari. These failures are surprising because discrete Q-networks lack the action-input distribution mismatch...
Daniel Palenicek, Mikael Henaff, Scott Fujimoto et al.· 0 citations
Evidence that parameter-space exploration can improve reinforcement learning for LLMs is presented, and a family of methods called Perturbed Parameter Policy Optimization (3PO) is introduced which use different sampling strategies and different rollout grouping for reward estimation.
Reinforcement learning (RL) has become central to post-training of large language models. Recent advances in RL for Mixture-of-Experts (MoE) models have primarily focused on improving optimization stability and training efficiency, while treating the expert selection as a fixed component. Since routing determines the s...
Hong-Yi He, Zheng-Wen Lin, Xiao Liu et al.· 0 citations
When rewards are sparse, reinforcement learning with verifiable rewards (RLVR) often uses hints or intermediate guidance to generate more successful rollouts. This enrichment biases policy-gradient updates unless corrected via importance weights, but existing methods omit correction or truncate importance weights in or...
Jinwoo Kim, Shraddha Barke· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.