Skip to content
Preprint

Momba: Network Modernization Improves Multi-Objective Reinforcement Learning

Aug 2026 · 0 citations · 43 references
Computer Science

TL;DR

Recent advances in neural network design are integrated: observation and feature normalization, weight normalization, and modeling of distributional returns with an entropy-regularized MORL algorithm, demonstrating that these changes substantially improve the quality of the produced solution sets without requiring major changes to the underlying algorithm.

Abstract

Recent advances in deep reinforcement learning (RL) have shown that improving neural network architectures can yield substantial gains in sample efficiency and asymptotic performance without altering the underlying algorithms. In contrast, work on multi-objective reinforcement learning (MORL), which aims to discover a set of policies that balance trade-offs among conflicting objectives, has predominantly focused on algorithmic innovations, leaving the area of architectures underexplored. While the optimal policies and value functions can differ significantly depending on the trade-offs, MORL algorithms commonly represent them with simple feedforward networks conditioned on the trade-off. This raises the question of whether the performance of the algorithms could be improved with more expressive function approximators. In this paper, we integrate recent advances in neural network design: (i) observation and feature normalization, (ii) weight normalization, and (iii) modeling of distributional returns with an entropy-regularized MORL algorithm. The empirical results across standard continuous control benchmarks demonstrate that these changes substantially improve the quality of the produced solution sets without requiring major changes to the underlying algorithm.

View source

Similar papers

Open access Oct 2026

Controllable Orthogonalization for Stabilizing Neural Network Training in Deep Reinforcement Learning

Deep reinforcement learning (DRL) relies on neural networks trained under temporally correlated experience, evolving replay distributions, and bootstrapped targets, conditions that can destabilize neural function approximation. This study investigates controllable orthogonalization as a network-level mechanism for impr...

Jia Yu, Dun Su, Hua Cui et al. · 0 citations
#machine learning Preprint Sep 2026

Towards Better Training Signal: Advantage Clipped Policy Optimization

Reinforcement learning (RL) has become a cornerstone for improving the reasoning capabilities of large language models (LLMs), but the need for on-policy data substantially limits training efficiency. Reusing off-policy data through importance sampling (IS) can improve efficiency but introduce considerable instability....

Rui-Chuan Huang, Jing-Han Liu, Cong-Liang Chen · 0 citations
#artificial intelligence Preprint Sep 2026

On BatchNorm Forward Modes in Value-Based Reinforcement Learning

Batch normalization (BN) substantially improves sample efficiency in continuous-control actor-critic methods such as CrossQ, yet recent studies report performance degradation in discrete-action value learning on Atari. These failures are surprising because discrete Q-networks lack the action-input distribution mismatch...

Daniel Palenicek, Mikael Henaff, Scott Fujimoto et al. · 0 citations
Preprint Aug 2026

Parameter Exploration for RLVR via Variational Learning

Evidence that parameter-space exploration can improve reinforcement learning for LLMs is presented, and a family of methods called Perturbed Parameter Policy Optimization (3PO) is introduced which use different sampling strategies and different rollout grouping for reward estimation.

Vatsal Venkatkrishna, Nico Daheim, Iryna Gurevych · 0 citations
#natural language process... Preprint Sep 2026

Expert-Space Exploration in MoE Reinforcement Learning

Reinforcement learning (RL) has become central to post-training of large language models. Recent advances in RL for Mixture-of-Experts (MoE) models have primarily focused on improving optimization stability and training efficiency, while treating the expert selection as a fixed component. Since routing determines the s...

Hong-Yi He, Zheng-Wen Lin, Xiao Liu et al. · 0 citations
#machine learning Preprint Oct 2026

Understanding Enrichment in Reinforcement Learning

When rewards are sparse, reinforcement learning with verifiable rewards (RLVR) often uses hints or intermediate guidance to generate more successful rollouts. This enrichment biases policy-gradient updates unless corrected via importance weights, but existing methods omit correction or truncate importance weights in or...

Jinwoo Kim, Shraddha Barke · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.