Skip to content

Sparse Gaussian-Mixture-Model Q-Functions via Hadamard Overparametrization for Online Reinforcement Learning

Jul 2026 · arXiv.org · Vol abs/2607.23474 · 0 citations · 55 references
Computer Science

TL;DR

This paper develops an online, off-policy policy-iteration framework for reinforcement learning (RL), based on sparse Gaussian-mixture-model Q-functions (S-GMM-QFs), enabling interpretable sparsification through smooth regularization that facilitates Riemannian-based optimization.

Abstract

This paper develops an online, off-policy policy-iteration framework for reinforcement learning (RL), based on sparse Gaussian-mixture-model Q-functions (S-GMM-QFs). The framework reconciles streaming, non-stationary data with the Riemannian structure of the parameter space while handling distributional mismatch through experience replay. S-GMM-QFs are introduced via Hadamard overparametrization, enabling interpretable sparsification through smooth regularization that facilitates Riemannian-based optimization. Overparametrization allows the framework to adaptively identify meaningful components from a large initial pool, yielding sparse models where interpretability emerges naturally from geometry: each component's parameters (means and covariances) explicitly encode its geometric role in the ambient state-action space. These geometric roles are learned through online gradient descent on a smooth objective over a (Cartesian-product) Riemannian manifold. Numerical tests demonstrate that S-GMM-QFs match or exceed deep RL methods while using substantially fewer parameters and achieving faster improvement per observed transition. Notably, parameter efficiency and interpretability combine to maintain strong generalization in low-parameter regimes where sparsified deep RL approaches degrade.

View source

Similar papers

Sparse Additive Off-Policy Evaluation for Reinforcement Learning with Potentially Limited Number of Trajectories

We develop a new framework for flexible, nonlinear, and interpretable off-policy evaluation for infinite-horizon reinforcement learning. To handle large state spaces and support transparent decision-making, we model the Q-function using a nonlinear function class with a sparse additive structure. We derive high-probabi...

Tuo-Yi Zhao, Cheng-Chun Shi, Zheng-Ling Qi et al. · 0 citations
Preprint Aug 2026

Information-Geometric Forward Policy Training in GFlowNets

Generative Flow Networks (GFlowNets) have emerged as a flexible framework for amortised inference over discrete and mixed discrete-continuous objects, requiring only an unnormalised target density specified through a reward. In this work, we formulate forward-policy training in GFlowNets through the information geometr...

Y. Raykov, Rodrigo Veiga · 1 citation
#machine learning Preprint Sep 2026

No-Regret Bayesian Optimization with Finite-Library Input-Warped Kernels

Gaussian-process Bayesian optimization (GP-BO) excels at black-box optimization of costly functions, e.g., hyperparameter optimization (HPO) and multi-agent system (MAS) design. Convergence-rate guarantees exist for select methods, notably GP upper confidence bound (GP-UCB), but require a fixed kernel. Critically, the...

Edvin Ketabati Augustinsson, Robert A. Bridges · 0 citations
#artificial intelligence Preprint Sep 2026

On BatchNorm Forward Modes in Value-Based Reinforcement Learning

Batch normalization (BN) substantially improves sample efficiency in continuous-control actor-critic methods such as CrossQ, yet recent studies report performance degradation in discrete-action value learning on Atari. These failures are surprising because discrete Q-networks lack the action-input distribution mismatch...

Daniel Palenicek, Mikael Henaff, Scott Fujimoto et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Noisy-Space Policy Gradient for Diffusion Policies in Offline Reinforcement Learning

Diffusion policies offer a powerful and expressive parameterization for continuous control. Yet, their integration with reinforcement learning remains conceptually and algorithmically challenging. In this work, we address this gap by introducing a noisy-space action-value (Q-)function that assigns values to diffusion l...

Mahmoud Selim, Cristina Cipriani, K. H. Johansson · 0 citations
Preprint Aug 2026

Revisiting TD Target Aggregation under Uncertainty in Q-Learning

The proposed SADQ is a simple modification to Q-learning that regularizes how the TD target is formed, and consistently improves training stability across classical control tasks, real-world vector-based environments, and Atari benchmarks when compared to strong DQN variants.

Li-Peng Zu, Xiaonan Zhang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.