Skip to content

Author

Wenwen Qiang

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Temporal GRPO: Beyond Trajectory-Level Credit in Vision-Language-Action Reinforcement Learning

Temporal GRPO addresses the problem of trajectory-level credit aliasing in post-train VLA policies by constructing detectable task stages, aligning each rollout with stage-specific action intervals, and comparing only rollouts that have entered the same stage.

Yao Zhou, Hang Gao, Fengge Wu et al. · 0 citations
#machine learning Preprint Aug 2026

GUPO: Gradient Uncertainty-aware Policy Optimization for Post-Training Large Language Models

Gradient Uncertainty-Aware Policy Optimization is proposed, which models each group gradient as a random variable under a Bayesian formulation and estimates its probability distribution and derives gradient uncertainty using a Dirichlet-based formulation and uses it to calibrate the contribution of each group gradient during aggregation.

Peizheng Guo, Jianqi Zhang, Xingyu Zhang et al. · 0 citations