Skip to content

Hardware Characterization of Di ! usion vs. Autoregressive Language Model Inference: Compute vs. Memory Bottlenecks

· 0 citations · 67 references

TL;DR

A hardware-level characterization of representative di ! usion language models is presented and it is shown that, although MDLMs and AR models share similar Transformer building blocks, di ! usion inference exhibits fundamentally different bottlenecks at the hardware level, breaking the assumptions underlying modern AR serving architectures and systems.

View source

Similar papers

Preprint Aug 2026

Serving Masked Diffusion LLMs: Characterization and Design Principles from Real Hardware

Results show that serving diffusion language models needs parallelism at the level of each denoising step, which differs from AR serving in how admission and eviction interact with an already shared forward pass.

Farhana Amin, Sabiha Afroz, Mona Moghadampanah et al. · 0 citations
#machine learning Preprint Sep 2026

Unlocking Lossless Speedups in LLMs via Discrete Diffusion

Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressive (AR) structure requires slow, sequential token generation. To overcome this bottleneck, we introduce diffusion-augmented LLMs, a new class of models that defines an AR model distribution while using diffusion...

S. Sahoo, Ling-Jie Chen, Khiem Pham et al. · 1 citation
Preprint Aug 2026

Pipeline-Native Transformers: Co-Designing Model Architecture and CPU Inference for Bandwidth-Efficient Autoregressive Decode

This report argues that the most effective response to single-token autoregressive decode on CPUs is to co-design the model architecture and the inference runtime together, and presents cflow, a CPU-first streaming engine, alongside a family of pipeline-native transformer architectures whose inter-layer dependency grap...

Tom Poperszky · 0 citations
Preprint Aug 2026

Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative Inference

Vision-language-action (VLA) models have emerged as a key component in embodied AI. Among existing approaches, diffusion-based VLA models achieve superior motion quality and generalization. However, diffusion-based VLA models are compute-intensive and must run at high control frequency, e.g., 50-200 Hz. Thus, it impose...

Zheng Liu, Zeyu Guo, Zihan Liu et al. · 1 citation
#natural language process... Preprint Sep 2026

Design of the IBM Granite 5.0 TurboCTC ASR Model

We describe the architecture, training methodology and inference speedups of Granite 5.0 Turbo CTC, a 470 million parameter encoder-only model with an excellent speed-accuracy tradeoff. The architecture uses pyramidal temporal subsampling within Conformer blocks using strided depthwise convolutions, block-diagonal (chu...

Brian Kingsbury, G. Saon, Masayuki Suzuki et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.