Skip to content

Author

Zhen-Xing Fan

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Hardware Characterization of Di ! usion vs. Autoregressive Language Model Inference: Compute vs. Memory Bottlenecks

A hardware-level characterization of representative di ! usion language models is presented and it is shown that, although MDLMs and AR models share similar Transformer building blocks, di ! usion inference exhibits fundamentally different bottlenecks at the hardware level, breaking the assumptions underlying modern AR...

Zhen-Xing Fan, Kevin Skadron · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.