Skip to content
Preprint

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning

Aug 2026 · 0 citations · 27 references
Computer Science

TL;DR

This work introduces Boundary-Seeking Policy Gradient (BSPG), a first-order method whose update combines a tangential component that improves reward while preserving cost to first order with a signed, residual-driven normal component that regulates the policy toward the active boundary from either side.

Abstract

Safe reinforcement learning maximizes reward subject to safety constraints. For Constrained Markov Decision Processes, the linear-programming view over occupancy measures implies that whenever the constraint is active at optimality, the optimal policy lies exactly on the constraint boundary, yet standard gradient-based methods do not exploit this structure and often settle in the feasible interior. We introduce Boundary-Seeking Policy Gradient (BSPG), a first-order method whose update combines a tangential component that improves reward while preserving cost to first order with a signed, residual-driven normal component that regulates the policy toward the active boundary from either side; the combined direction admits an algebraic Lagrangian form with an induced coefficient and no learned dual variable. Under exact gradients and stated regularity conditions, the constraint residual converges to zero from either side with a finite-horizon $O(1/\sqrt{T})$ bound, the tangential component is a reward-ascent direction on the boundary, and any convergent parameter sequence is stationary on the active constraint set, satisfying the KKT conditions when the limit is also a local maximizer over the feasible set. This complements existing analyses, which certify feasibility but do not characterize the constraint value at convergence. On a standard Safety-Gymnasium navigation task, BSPG attains higher reward while tracking the boundary more tightly than the compared baselines.

View source

Similar papers

Preprint Sep 2026

SafePG: Safe and Globally Optimal Reinforcement Learning with Hard Constraints

We present an optimal and convergent model-free policy gradient (PG) reinforcement learning (RL) framework for controlling nonlinear dynamical systems under hard safety constraints. We first construct a class of stochastic wrapper policies centered around a deterministic controller, thereby enabling exploration in unkn...

Vipul K. Sharma, Wesley A. Suttle, S. Sivaranjani · 0 citations
Preprint Aug 2026

On the Optimality of Markovian Policies for Chance-Constrained Covariance Steering

It is shown that, even for the covariance steering problem with a broad class of commonly used state and control safety constraints, the synthesized Markovian policy almost surely produces the same control actions as the history-dependent policy and therefore the same state trajectories, cost, and moments.

Naoya Kumagai, K. Oguri · 0 citations
Preprint Aug 2026

Hidden Star-Convexity in Policy Optimization for Gain-Scheduled LQR: Extended Version

We study policy optimization for gain-scheduled linear quadratic regulation, where one schedule of gains, interpolated through fixed weighting functions, is optimized against a family of plants. The resulting cost can develop spurious local minima, and existing convergence certificates are either local or severely cons...

Shiva Shakeri, Péter Baranyi, M. Mesbahi · 0 citations
Preprint Aug 2026

Learning-Based Stochastic Optimal Control with Infinite-Horizon Probabilistic Constraints

In this paper, we consider stochastic optimal control problems with infinite-horizon joint chance constraints. By means of an appropriate state augmentation, we reformulate the original problem as a constrained Markov decision process, in which both the cost and the constraint function exhibit an additive structure. We...

Francesco Cordiano, Kang-Hui He, B. de Schutter · 0 citations
#machine learning Preprint Sep 2026

Local and Global Stability in Performative Reinforcement Learning

This work separates local mixed stability, an occupancy-weighted first-order relaxation that is equivalent to stationarity, from global mixed stability, which certifies against arbitrary deviating policies, and extends both notions to n-player performative Markov games, obtaining local stability with no assumption on t...

Debmalya Mandal · 0 citations
Preprint Sep 2026

Parallel Policy-Gradient Methods for Parameter Optimization of Nonlinear Feedback Controllers

Structured feedback controllers provide rigorous stability guarantees, but often require manual parameter tuning to achieve good closed-loop performance. Policy-gradient methods offer a systematic approach to parameter optimization; however, conventional gradient evaluation requires sequential forward state rollout and...

A. Nguyen, Leilei Cui · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.