LeanGRPO: Eliminating Redundant Recomputation in Diffusion RL
LeanGRPO is presented by restructuring the data-parallel layout and introducing two recompute-free training schedules for trajectory-logprob diffusion RL, which achieves up to 1.83x end-to-end speedup while preserving the original optimization objective.