Skip to content

Induction in Both Directions: A Mechanistic Analysis of In-Context Learning in Masked Diffusion Language Models

Jul 2026 · arXiv.org · Vol abs/2607.15893 · 0 citations · 38 references
Computer Science

TL;DR

This work studies how diffusion language models implement induction, a mechanism behind in-context learning in which the model finds a repeated context and copies the token that followed it, and compares attention-only AR models and absorbing-mask DLMs with matched architectures.

Abstract

While the internal mechanisms of autoregressive (AR) transformers have been studied extensively, much less is known about diffusion language models (DLMs), an emerging alternative that generates text by iterative denoising. In this work, we study how DLMs implement induction, a mechanism behind in-context learning in which the model finds a repeated context and copies the token that followed it. Our analysis compares attention-only AR models and absorbing-mask DLMs with matched architectures. We find that DLMs learn a bidirectional induction circuit, where previous-token and next-token heads write local context into the residual stream and later induction heads use it to find and copy the answer from the matching source position. The circuit is direction-symmetric, working whether the source appears in the past or in the future. When only left context is visible, matching what an AR model sees, the DLM does not outperform its AR counterpart in induction capabilities. However, we observe it has stronger induction when both sides of the masked token are visible, pointing to bidirectional context access rather than a stronger one-sided mechanism. Beyond induction, we provide causal evidence that DLMs compute the global fraction of masked tokens and use it as an implicit timestep, even though they are given no explicit timestep embedding.

View source

Similar papers

#natural language process... Preprint Sep 2026

DA-DLM: Explicitly Modeling Token Dependencies in Diffusion Language Models

Diffusion Language Models (DLMs) generate text by iteratively denoising a masked sequence, independently predicting multiple tokens at each step. This conditional independence discards inter-token dependencies and degrades coherence-an issue that parallels the multi-modality problem in Non-Autoregressive Translation (N...

Peng-Yu Ji, Zi-Chen Zhang, Xiang Hu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Zarya: A Hybrid Autoregressive--Masked Diffusion Language Model with Flexible Training and Dual-Mode Inference

Autoregressive language models (ARMs) are constrained by sequential, left-to-right generation, while masked diffusion models (MDMs) enable parallel decoding but suffer from high computational overhead due to the inability to reuse Key-Value (KV) cache and from incoherent generation arising from learning dependencies ov...

Leonid S. Sinev, Ilya Koziev, Vladislav Leshchuk · 0 citations
Jul 2026

PreDiff-LM: Pretrained Discrete Masked Diffusion Language Modeling with Hybrid Attention

PreDiff-LM preserves causal attention within the observed prompt while allowing full bidirectional attention within the masked target, position hybrid attention as a complementary mechanism for adapting pretrained causal backbones, while making explicit the remaining quality and inference-efficiency gaps to optimized A...

Zhengtao Yao, Run-Hao Li, Xu-Peng Chen et al. · 0 citations
#machine learning Preprint Sep 2026

dQwen3.5: Hybrid-Attention Diffusion Language Models

Adapting a pretrained autoregressive (AR) model is a cost-efficient route to a diffusion language model (DLM). While nearly all such adaptations start from a full-attention transformer, AR modeling has shifted toward hybrid architectures that interleave attention and RNN layers. This creates an obstacle for adaptation:...

Anton Xue, Litu Rout, Aditya Akella et al. · 0 citations
#artificial intelligence Preprint Sep 2026

In-Place Instruction Following in Diffusion Language Models

Diffusion Large Language Models (dLLMs) generate text via bidirectional iterative denoising, naturally supporting user-specified constraints anchored at arbitrary output positions, a paradigm known as In-place Prompting (IPP). We formalize this as the In-place Instruction Following (IIF) task and construct IIF-Bench, a...

Zheng Nie, Zhe-Rui Li, Jia-Ming Zhang et al. · 0 citations
Jul 2026

Beyond the Bidirectional Promise: Re-evaluating the Robustness of Diffusion Language Models

While highly stochastic DLM loss landscapes naturally resist gradient-based adversarial suffixes, they provide no guaranteed defense against natural noise, proving that everyday robustness is weight-dependent rather than inherently architectural.

Saurabh Yadav, B. Patro, V. Agneeswaran · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.