It is shown that, even for the covariance steering problem with a broad class of commonly used state and control safety constraints, the synthesized Markovian policy almost surely produces the same control actions as the history-dependent policy and therefore the same state trajectories, cost, and moments.
Abstract
Many studies on finite-horizon stochastic optimal control, including covariance steering, parameterize control policies as state-history-affine. This parameterization enables a convex reformulation, thereby yielding a tractable solution method. However, the necessity of dependence on previous states has not been well established. \textit{Is this dependence necessary, or merely an artifact of the convex reformulation?} We show that it is an artifact that can be removed losslessly. Given an optimal solution of the state-history-affine formulation, we construct a deterministic Markovian policy which is affine in the current state. We show that, even for the covariance steering problem with a broad class of commonly used state and control safety constraints, the synthesized Markovian policy almost surely produces the same control actions as the history-dependent policy and therefore the same state trajectories, cost, and moments. Thus, every optimum of the history-dependent formulation admits a lossless Markovian transformation. Geometrically, the history-dependent formulation lifts the policy space for convexity, and its optimal solution can be projected back to the Markovian policy space. We extend the analysis to output feedback and a convex upper-bounding surrogate for value-at-risk costs.
In this paper, we consider stochastic optimal control problems with infinite-horizon joint chance constraints. By means of an appropriate state augmentation, we reformulate the original problem as a constrained Markov decision process, in which both the cost and the constraint function exhibit an additive structure. We...
Francesco Cordiano, Kang-Hui He, B. de Schutter· 0 citations
This paper provides the first finite-time convergence guarantees for this algorithm in this setting, for which it is proved that NPG converges sublinearly with a rate of $\mathcal{O}(H^{2}/t)$ after $t$ iterations, where $H$ is the horizon length.
Asha Barua, S. Khodadadian· arXiv.org· 0 citations
We study optimal input design over a finite horizon for linear dynamical systems. The goal is to minimize a weighted inverse-covariance (information) criterion subject to an energy budget. The set of covariances achievable by causal policies is convex but lacks a tractable explicit description, ruling out projection-ba...
Fethi Bencherki, Bruce D. Lee, Nikolai Matni et al.· 0 citations
This work introduces Boundary-Seeking Policy Gradient (BSPG), a first-order method whose update combines a tangential component that improves reward while preserving cost to first order with a signed, residual-driven normal component that regulates the policy toward the active boundary from either side.
Chenhua Fan, Jiahui Zhu, Yuhang Zhang et al.· 0 citations
We study policy optimization for gain-scheduled linear quadratic regulation, where one schedule of gains, interpolated through fixed weighting functions, is optimized against a family of plants. The resulting cost can develop spurious local minima, and existing convergence certificates are either local or severely cons...
Shiva Shakeri, Péter Baranyi, M. Mesbahi· 0 citations
A novel spectrum assignment method is proposed to obtain an initial stabilizer for PI in continuous-time indefinite stochastic linear quadratic control with the help of the Lyapunov-type operator's spectrum, which is gradually approximated from the stable auxiliary system by adjusting a cumulative factor, thereby obtai...
Xinyu Cao, Bing-Chang Wang, Ying Cao· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.