Skip to content

Reinforcing the Generation Order of Multimodal Masked Diffusion Models

Jul 2026 · arXiv.org · Vol abs/2607.08056 · 0 citations · 40 references
Computer Science Mathematics

TL;DR

This work introduces a learnable control module trained via Group Relative Policy Optimization (GRPO) to determine the generation order and demonstrates that learning this control block substantially improves both text-to-image alignment and multimodal understanding in DLMs.

Abstract

Diffusion Language Models (DLMs) have recently achieved substantial progress in natural language generation tasks. Recent research demonstrates that adaptive token generation ordering can significantly improve performance in mathematical reasoning and code synthesis applications. In this work, we investigate the optimization of generation order for both text-to-image synthesis and multimodal understanding. We first establish that, unlike structured problems in language generation such as Sudoku puzzles, model logits alone are insufficient for determining optimal generation sequences in text-to-image generation and multimodal understanding. To address this challenge, we introduce a learnable control module trained via Group Relative Policy Optimization (GRPO) to determine the generation order. Our results demonstrate that learning this control block substantially improves both text-to-image alignment and multimodal understanding in DLMs. In particular, it enhances the model's ability to capture fine-grained spatial relationships in generated images while also strengthening performance on multimodal reasoning and comprehension tasks. We evaluate our framework on GenEval, an object-focused benchmark for text-to-image alignment, where it achieves 4.08% relative improvements. In addition, experiments on VLMEvalKit confirm 4.85% relative improvements in multimodal understanding, highlighting the broad effectiveness of our approach.

View source

Similar papers

Preprint Aug 2026

Visual Information-Guided Parallel Decoding for Diffusion Multimodal Large Language Models

The Visual Information-Guided Sampler (VIG-Sampler), which prioritizes tokens based on their attention to image tokens, and imposes a constraint that penalizes candidate tokens whose image-attention distributions are similar to those of previously selected tokens, thereby increasing the information gain of the decoded...

Insu Lee, Woo-Soon Park, Wonseok Shin et al. · 1 citation
#natural language process... Preprint Sep 2026

Representation-based Masked Diffusion Model

Masked Diffusion Models (MDMs) have emerged as a compelling paradigm for language modeling, offering the capability for efficient parallel text generation. However, existing parallel sampling methods typically update multiple masked tokens independently and ignore the complex mutual dependencies among the masked tokens...

Yang Hu, Ding Huang, Xue-Yu Zhou et al. · 0 citations
#artificial intelligence Preprint Sep 2026

In-Place Instruction Following in Diffusion Language Models

Diffusion Large Language Models (dLLMs) generate text via bidirectional iterative denoising, naturally supporting user-specified constraints anchored at arbitrary output positions, a paradigm known as In-place Prompting (IPP). We formalize this as the In-place Instruction Following (IIF) task and construct IIF-Bench, a...

Zheng Nie, Zhe-Rui Li, Jia-Ming Zhang et al. · 0 citations
#machine learning Preprint Sep 2026

The Alignment Illusion in Multimodal Large Language Models

Internal visual-text alignment in MLLMs is best read as a geometric diagnostic of the visual stream inside the language model rather than a direct proxy for content-level cross-modal interaction, and is most informative when calibrated by controlled task evidence.

Hong-Han Wang, Yun-Tao Wang, Hui-Chao Ding · 0 citations
#machine learning Preprint Sep 2026

ELF-REG: Scaling Continuous Diffusion Language Models to Reasoning Tasks

This work introduces ELF-REG, which improves learning with representation alignment and entanglement (REPA+REG), where a frozen AR teacher supervises intermediate denoiser features and supplies a global representation that is jointly denoised with the response.

Z. Li, William Chen, Bing-Shuo Qian et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.