Skip to content

Author

Pei-Lin Zhao

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

Matrix AdaGrad: Row-wise and Column-wise Adaptive Subgradient Methods

An Online Mirror Descent framework with adaptive proximal functions for matrix-valued parameters, providing a principled methodology for deriving matrix-aware adaptive optimization through online regret minimization and yields Row-wise Matrix AdaGrad and Column-wise Matrix AdaGrad as concrete instantiations with regret...

Wenpeng Zhang, Run-Sheng Yu, Pei-Lin Zhao · 0 citations
#artificial intelligence Preprint Sep 2026

Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening

Value Flattening is identified as an important yet overlooked failure mode of critic learning in standard PPO and a simple sparse supervision strategy can mitigate it; SParse Proximal Policy Optimization is introduced, which applies the value loss to only a few well-separated states in each response to mitigate both ef...

Yi-Zhuo Li, Jian-Hao Yan, Yun Luo et al. · 1 citation
#artificial intelligence Preprint Sep 2026

T-LoopFormer: Token-Level Elastic-Depth Looped Transformers for Latent Reasoning with Dynamic Routing

This work proposes dynamic token-choice routing for looped transformers, enabling each token to adaptively determine its own number of loop iterations based on its hidden state, which can improve the token generation accuracy and validate the effectiveness of token-choice router and recursion-wise KV cache.

Ming-Qian Yu, Wen-Peng Zhang, Shao-Bo Cui et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.