Reinforcing the Generation Order of Multimodal Masked Diffusion Models
This work introduces a learnable control module trained via Group Relative Policy Optimization (GRPO) to determine the generation order and demonstrates that learning this control block substantially improves both text-to-image alignment and multimodal understanding in DLMs.