LeanGRPO: Eliminating Redundant Recomputation in Diffusion RL
LeanGRPO is presented by restructuring the data-parallel layout and introducing two recompute-free training schedules for trajectory-logprob diffusion RL, which achieves up to 1.83x end-to-end speedup while preserving the original optimization objective.
Si-Jie Wang, Zhi-Qiang Tan, Xin-Rui Yang et al.
· 0 citations