#machine learning
Jun 2025
Any-Order GPT as Masked Diffusion Model: Decoupling Formulation and Architecture
It is shown decoder-only MDMs, despite a larger modeling space, can achieve significant inference speedups and comparable perplexity with techniques like temperature annealing, offering a path to reduced inference compute.
Shuchen Xue, Tianyu Xie, Tianyang Hu et al.
· arXiv.org · 17 citations
· ⚡1