Skip to content

Author

Qiuyang Mang

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

EasyPPO: Stabilizing the Critic Is Key

A key strength of Proximal Policy Optimization (PPO) is its learned critic, which uses historical trajectories collected during reinforcement learning to estimate expected returns and reduce policy-gradient variance. However, we find that the critic is also a major source of instability in reinforcement learning for la...

Xuan-Yi Zhou, Qiu-Yang Mang, Huan-Zhi Mao et al. · 0 citations
#natural language process... Preprint Sep 2026

When Agents Slow Down: Understanding LLM Agents'Test-Time Strategies via Elo-per-token Analysis

This work defines the scaling inflection point as the per-session budget where marginal Elo gains match the independent-sampling reference, and proposes Elo-per-token analysis, which tracks the best solution found at each token budget and uses a Bradley-Terry model to aggregate within-task orderings into Elo ratings ac...

Kai-Yuan Liu, Qiu-Yang Mang, Bo-Fei Peng et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.