Skip to content
Preprint

Ranking-Augmented On-Policy Optimization with Adaptive Advantage-Normalization for Constrained Control

Aug 2026 · 1 citation · 37 references
Engineering Computer Science

Abstract

This paper analyzes the boundedness and feasibility properties of Advantage-Ranked Group Relative Policy Optimization (A-GRPO), a ranking-augmented, critic-free policy gradient method employing a Transformer-encoder actor for fixed-horizon control with terminal constraints. When feasibility is evaluated only at the final step, the resulting sparse feedback destabilizes critic-based advantage estimation and weakens standard Lagrangian approaches. A trajectory-level ranking mechanism that augments group-relative policy updates by reweighting advantages according to constraint satisfaction is formalized, and three results are established: (i) a scale-adaptive per-timestep normalization bounds advantage variance at every timestep independently, (ii) the ranked advantage strictly separates feasible from violating trajectories under a verifiable ranking-weight condition, biasing the policy gradient toward constraint satisfaction, and (iii) the adaptive dual variables remain bounded and exhibit a drift-balance property that acts as a feedback mechanism for feasibility. These results are validated on a 3,605-step series-hybrid powertrain energy management task with a terminal state-of-charge constraint, where A-GRPO achieves 75.4% mean sustained feasibility with return within 3.7% of the dynamic programming optimum, outperforming a Proximal Policy Optimization with Lagrangian penalties (PPO-Lag) baseline (27.4% sustained), and ablation experiments confirm that both the ranking and Lagrangian components are necessary for this performance.

View source

Similar papers

Preprint Aug 2026

Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization

Environment-Regularized Policy Optimization (ERPO) replaces the standard Policy-KL regularizer while achieving effective control over query distribution drift, delivering stronger accuracy and substantially more stable behavior under high-temperature decoding and long-horizon training.

Xianlei Zhou, Xiangdi Meng, Yu He et al. · 0 citations
#artificial intelligence Preprint Sep 2026

GVPO++: Group Variance Policy Optimization for LLM Post-Training and On-Policy Distillation

Post-training plays a pivotal role in enhancing the reasoning capabilities and task-specific expertise of large language models (LLMs). Despite recent advances in post-training methods, such as Group Relative Policy Optimization (GRPO), their practical deployment remains impeded by training instability arising from the...

Kai-Chen Zhang, Yuzhong Hong, Jun-Wei Bao et al. · 0 citations
Preprint Aug 2026

Hidden Star-Convexity in Policy Optimization for Gain-Scheduled LQR: Extended Version

We study policy optimization for gain-scheduled linear quadratic regulation, where one schedule of gains, interpolated through fixed weighting functions, is optimized against a family of plants. The resulting cost can develop spurious local minima, and existing convergence certificates are either local or severely cons...

Shiva Shakeri, Péter Baranyi, M. Mesbahi · 0 citations
Preprint Aug 2026

Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization

Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization (SA-MRPO), which standardizes each reward objective independently and adaptively discounts its contribution according to a batch-level estimate of objective saturation, and dynamically reallocates optimization effort toward under-optimized obje...

Yi-Xuan Wang, Yi-Fei Chen, Haichao Zhang et al. · 0 citations
Preprint Sep 2026

Critic-Free Policy Iteration for Continuous-Time Zero-Sum Games: A Policy-Space Riccati Approach

This paper develops a critic-free policy iteration (PI) method for continuous-time linear zero-sum games. The central idea is to characterize the saddle-point policies directly in the joint policy space, rather than treating the quadratic value matrix as an iterative variable. A policy game Riccati equation (PGRE) is i...

Jia-Cheng Wu, Yang Zhu, Hong-Ye Su · 0 citations
Preprint Sep 2026

Sparse One-Step-Ahead Optimal Control of Time-Varying Affine Opinion Networks: Tracking and Competitive Games

This paper studies resource-limited external influence in time-varying opinion networks when a controller must choose both a small set of agents and a scalar intervention at each update. We use the affine free response, which includes DeGroot and Friedkin--Johnsen dynamics, followed by a direct sparse action. Eliminati...

G. Gentil, Amit Bhaya · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.