Skip to content

Author

Changwen Zheng

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Aug 2026

GUPO: Gradient Uncertainty-aware Policy Optimization for Post-Training Large Language Models

Gradient Uncertainty-Aware Policy Optimization is proposed, which models each group gradient as a random variable under a Bayesian formulation and estimates its probability distribution and derives gradient uncertainty using a Dirichlet-based formulation and uses it to calibrate the contribution of each group gradient during aggregation.

Peizheng Guo, Jianqi Zhang, Xingyu Zhang et al. · 0 citations