Skip to content

Scalable and Robust Reinforcement Learning Through Expectile Regression.

Aug 2026 · IEEE Transactions on Neural Networks and Learning Systems · Vol PP, pp. 1-8 · 0 citations
Medicine

Abstract

Robust reinforcement learning (RRL) aims to develop a robust policy that maintains stable performance across diverse environments characterized by an uncertainty set. This set consists of perturbed environments derived from a nominal (training) environment that generates samples, thereby capturing potential discrepancies between training and real-world conditions. Recently, an adjacent uncertainty set has been introduced, providing more realistic perturbations compared to conventional formulations. Despite its solid theoretical foundation, the existing sample-based implementation of the robust Bellman update suffers from limited scalability and practical applicability in real-world scenarios. In this brief, we present, for the first time, scalable RRL algorithms that overcome these challenges by leveraging expectile regression. Extensive experiments demonstrate that the proposed methods significantly enhance the robustness of state-of-the-art (SOTA) RL algorithms while maintaining a practical computational cost comparable to strong off-policy baselines. In particular, the proposed methods exhibit up to a 23.6% average improvement in robustness under environmental perturbations over SOTA RL baselines while maintaining comparable computational complexity.

View source

Similar papers

Forecasting in Offline Reinforcement Learning for Non-stationary Environments

F orecasting in Non-stationary O ffline RL (F ORL), a framework that combines zero-shot forecasting with the agent’s experience, aims to bridge the gap between offline RL and the complexities of real-world, non-stationary environments.

Suzan Ece, Georg Martius, Emre Ugur et al. · 0 citations
#machine learning Preprint Sep 2026

Local and Global Stability in Performative Reinforcement Learning

This work separates local mixed stability, an occupancy-weighted first-order relaxation that is equivalent to stationarity, from global mixed stability, which certifies against arbitrary deviating policies, and extends both notions to n-player performative Markov games, obtaining local stability with no assumption on t...

Debmalya Mandal · 0 citations
#artificial intelligence Preprint Sep 2026

SUN: Reaching for Novelty in Reinforcement Learning

This paper proposes SUccessor-to-Novelty (SUN), an indicator derived from successor value functions to identify goals that are both novel and reachable and presents an adaptive goal-selection strategy that leverages these properties, and an accurate yet lightweight pseudocount to avoid the overhead of classic methods.

Wen-Yan Yang, A. Mustafin, Dominik Baumann et al. · 0 citations
Preprint Sep 2026

Policy Gradient over History-Dependent Policy Classes for LQR with Domain Randomization

Domain Randomization (DR) has been widely used to overcome the sim-to-real gap by training a controller on a distribution of simulated environments via reinforcement learning. While DR can achieve robust performance simply using controllers synthesized via policy gradient (PG) methods, the optimization landscape is not...

Tesshu Fujinami, Bruce D. Lee, Anastasios Tsiamis et al. · 0 citations
#machine learning Preprint Sep 2026

Model-based Bootstrap for Offline Policy Evaluation in Tabular Reinforcement Learning

This paper proposes a model-based bootstrap framework for uncertainty quantification of OPE in finite-horizon, time-inhomogeneous Markov decision processes (MDPs), and establishes bootstrap distributional consistency, asymptotically valid confidence intervals, and consistent variance estimation for the target policy va...

Wei-Wei Wang, Yu-Qiang Li, Xian-Yi Wu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Learning Perturbation Robust Policies for LLM Agents with Stable Optimization

Reinforcement learning (RL) has become an effective post-training paradigm for long-horizon large language model (LLM) agents. However, we find that the resulting policies can be sensitive to various policy perturbations, such as hidden-state noise, pruning, and quantization. In this work, we study how to improve pertu...

Peng-Xin Wang, Yuan-Zhe Li, Yuxin Ren et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.