Skip to content

Author

George J. Pappas

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

Policy Gradient over History-Dependent Policy Classes for LQR with Domain Randomization

Domain Randomization (DR) has been widely used to overcome the sim-to-real gap by training a controller on a distribution of simulated environments via reinforcement learning. While DR can achieve robust performance simply using controllers synthesized via policy gradient (PG) methods, the optimization landscape is not well understood, even in the case of linear quadratic regulator (LQR) objectives. To this end, we first study PG of domain randomized LQR over history-dependent policy classes, such as finite impulse response controllers, as they can extend the possibilities of simultaneous stabilization. Second, to find such a stabilizing controller, we propose a curriculum learning based algorithm which gradually expands the memory of the controller. Finally, we show that PG with the proposed algorithm converges globally to the minimizer of a sample average approximation of the DR objective under suitable bounds on the heterogeneity of environments. Empirical results support our findings and highlight promising directions for future work, including nonlinear domain-randomized control.

Tesshu Fujinami, Bruce D. Lee, Anastasios Tsiamis et al. · 0 citations
Preprint Sep 2026

Verifying performance, stability, and feasibility of inexact non-linear model predictive controllers

We introduce a verification framework to numerically analyze inexact model predictive controllers (MPCs) in the constrained non-linear discrete-time setting. Rather than modifying the controller so that guarantees hold by construction, we treat the controller as given. In particular, we focus on two types of inexact controllers: (a) one whose input is extracted from a primal-dual point satisfying the Karush-Kuhn-Tucker (KKT) conditions of the non-convex MPC problem, and (b) one whose input is obtained by linearizing the dynamics and solving a convex quadratic program. The main idea of our verification framework is to formulate an optimization problem that searches over the worst-case initial state within a given set and control inputs consistent with the inexact controller to maximize a carefully-chosen performance metric. Using this framework, we show how to certify (i) the worst-case suboptimality gap of a single MPC problem, (ii) the worst-case closed-loop suboptimality gap over a given number of dynamical system iterations, (iii) closed-loop stability, and (iv) feasibility of the closed-loop system. Through numerical examples, we showcase the ability of our framework to precisely quantify both types of suboptimality, and to test the stability and feasibility of the inexact controllers.

Rajiv Sambharya, S. C. Anand, George J. Pappas · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.