Hardware Characterization of Di ! usion vs. Autoregressive Language Model Inference: Compute vs. Memory Bottlenecks
A hardware-level characterization of representative di ! usion language models is presented and it is shown that, although MDLMs and AR models share similar Transformer building blocks, di ! usion inference exhibits fundamentally different bottlenecks at the hardware level, breaking the assumptions underlying modern AR...