This work introduces Boundary-Seeking Policy Gradient (BSPG), a first-order method whose update combines a tangential component that improves reward while preserving cost to first order with a signed, residual-driven normal component that regulates the policy toward the active boundary from either side.
Abstract
Safe reinforcement learning maximizes reward subject to safety constraints. For Constrained Markov Decision Processes, the linear-programming view over occupancy measures implies that whenever the constraint is active at optimality, the optimal policy lies exactly on the constraint boundary, yet standard gradient-based methods do not exploit this structure and often settle in the feasible interior. We introduce Boundary-Seeking Policy Gradient (BSPG), a first-order method whose update combines a tangential component that improves reward while preserving cost to first order with a signed, residual-driven normal component that regulates the policy toward the active boundary from either side; the combined direction admits an algebraic Lagrangian form with an induced coefficient and no learned dual variable. Under exact gradients and stated regularity conditions, the constraint residual converges to zero from either side with a finite-horizon $O(1/\sqrt{T})$ bound, the tangential component is a reward-ascent direction on the boundary, and any convergent parameter sequence is stationary on the active constraint set, satisfying the KKT conditions when the limit is also a local maximizer over the feasible set. This complements existing analyses, which certify feasibility but do not characterize the constraint value at convergence. On a standard Safety-Gymnasium navigation task, BSPG attains higher reward while tracking the boundary more tightly than the compared baselines.
We present an optimal and convergent model-free policy gradient (PG) reinforcement learning (RL) framework for controlling nonlinear dynamical systems under hard safety constraints. We first construct a class of stochastic wrapper policies centered around a deterministic controller, thereby enabling exploration in unkn...
Vipul K. Sharma, Wesley A. Suttle, S. Sivaranjani· 0 citations
It is shown that, even for the covariance steering problem with a broad class of commonly used state and control safety constraints, the synthesized Markovian policy almost surely produces the same control actions as the history-dependent policy and therefore the same state trajectories, cost, and moments.
We study policy optimization for gain-scheduled linear quadratic regulation, where one schedule of gains, interpolated through fixed weighting functions, is optimized against a family of plants. The resulting cost can develop spurious local minima, and existing convergence certificates are either local or severely cons...
Shiva Shakeri, Péter Baranyi, M. Mesbahi· 0 citations
In this paper, we consider stochastic optimal control problems with infinite-horizon joint chance constraints. By means of an appropriate state augmentation, we reformulate the original problem as a constrained Markov decision process, in which both the cost and the constraint function exhibit an additive structure. We...
Francesco Cordiano, Kang-Hui He, B. de Schutter· 0 citations
This work separates local mixed stability, an occupancy-weighted first-order relaxation that is equivalent to stationarity, from global mixed stability, which certifies against arbitrary deviating policies, and extends both notions to n-player performative Markov games, obtaining local stability with no assumption on t...
Structured feedback controllers provide rigorous stability guarantees, but often require manual parameter tuning to achieve good closed-loop performance. Policy-gradient methods offer a systematic approach to parameter optimization; however, conventional gradient evaluation requires sequential forward state rollout and...
A. Nguyen, Leilei Cui· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.