Skip to content

Algorithmic Optimality Guarantees for Nonsmooth $H_\infty$ Output-Feedback Policy Search

Sep 2026 · 0 citations · 26 references
Mathematics Computer Science Engineering

TL;DR

It is proved that on the exact identity-gauge slice of the extended convex lift, $\varepsilon$-stationarity yields $O(\varepsilon)$-suboptimality on compact exact slices, which in turn yields convergence-rate guarantees for nonsmooth policy-search methods.

Abstract

We study continuous-time full-order dynamic output-feedback $H_\infty$ policy search, a nonconvex and nonsmooth problem. Direct policy search is a central paradigm in reinforcement learning and continuous control, but rigorous guarantees remain scarce in robust output-feedback settings. The $H_\infty$ problem is a canonical benchmark because it captures disturbance attenuation and robustness while exposing the hard nonsmooth geometry of policy-space optimization. We prove that on the exact identity-gauge slice of the extended convex lift, $\varepsilon$-stationarity yields $O(\varepsilon)$-suboptimality on compact exact slices, which in turn yields convergence-rate guarantees for nonsmooth policy-search methods. This result addresses the finite-time optimality-gap question raised by Guo and Hu [2022] in the more general dynamic output-feedback $H_\infty$ policy-search setting. We further use the established value equivalence supplied by extended convex lifting to formulate a nonstrict-feasibility bisection method with one final strict-feasibility recovery step, yielding an explicit $\varepsilon$-optimal stabilizing controller. These results provide a quantitative and algorithmic strengthening of prior qualitative optimality theory for nonsmooth $H_\infty$ policy search.

View source

Similar papers

Preprint Sep 2026

Weak Convexity and Proximal Bundle Methods for Nonsmooth Policy Optimization in Robust Control

We study policy optimization for discrete-time robust $\mathcal{H}_\infty$ control with static output-feedback, and present the first feasibility-preserving algorithm with a deterministic, non-asymptotic complexity guarantee. This problem naturally leads to a nonsmooth and nonconvex optimization over the set of stabili...

Yuto Watanabe, Feng Liao, Yang Zheng · 1 citation
Preprint Aug 2026

Zeroth-Order Nonsmooth Nonconvex Optimization with Convex Liftings and Its Application to State-Feedback $H_\infty$ Policy Optimization

This work proposes a zeroth-order proximal point algorithm and verifies that the assumptions underlying the analysis hold for discrete-time state-feedback state-feedback policy optimization, yielding an oracle complexity of $\widetilde{O}\left(n_u n_x\epsilon^{-3}\right)$ for attaining a prescribed objective value gap.

Xuhao Wang, Yu-Jie Tang · 2 citations · ⚡1
Preprint Sep 2026

Residual Feedback for Transformed-State Equilibrium Seeking via General Variational Inequalities

Motivated by networked systems in which equilibrium conditions apply to a regulated state rather than directly to the decision variable, we study transformed-state general variational inequalities. Using a projection-residual reformulation, we develop residual feedback methods that operate in the decision space without...

Griffin Smith, Afrooz Jalilzadeh · 0 citations
Preprint Sep 2026

A Riccati Approach to Mixed $H_2/H_\infty$ Closed-Loop Games for Infinite-Dimensional Stochastic Systems

This paper studies a finite-horizon mixed $H_2/H_\infty$ feedback Nash game for stochastic evolution equations on a separable Hilbert space. The drift generator is unbounded, the remaining coefficients are bounded, and the one-dimensional Brownian diffusion depends on the state, control, and disturbance. The $H_2$ chan...

Ming-Yang Shen, Wei-Hai Zhang, Qing-Xin Meng et al. · 0 citations
#machine learning Preprint Sep 2026

Sharp Critical Minimax Laws and No-Learning Thresholds in Continuous-Time Adaptive Control

We study episodic continuous-time control with an unknown vector control gain, scalar state, quadratic action cost, and smooth convex terminal cost. In the scalar Gaussian experiment, let $\Delta(H)$ denote the minimax improvement over zero control and set $\delta=\sqrt2H^2-1$. We prove the critical law $$ \Delta(H)\as...

Jia Chen · 0 citations
Preprint Sep 2026

Optimal input design via Frank-Wolfe

We study optimal input design over a finite horizon for linear dynamical systems. The goal is to minimize a weighted inverse-covariance (information) criterion subject to an energy budget. The set of covariances achievable by causal policies is convex but lacks a tractable explicit description, ruling out projection-ba...

Fethi Bencherki, Bruce D. Lee, Nikolai Matni et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.