A hardware-level characterization of representative di ! usion language models is presented and it is shown that, although MDLMs and AR models share similar Transformer building blocks, di ! usion inference exhibits fundamentally different bottlenecks at the hardware level, breaking the assumptions underlying modern AR serving architectures and systems.
Results show that serving diffusion language models needs parallelism at the level of each denoising step, which differs from AR serving in how admission and eviction interact with an already shared forward pass.
Farhana Amin, Sabiha Afroz, Mona Moghadampanah et al.· 0 citations
Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressive (AR) structure requires slow, sequential token generation. To overcome this bottleneck, we introduce diffusion-augmented LLMs, a new class of models that defines an AR model distribution while using diffusion...
S. Sahoo, Ling-Jie Chen, Khiem Pham et al.· 1 citation
This report argues that the most effective response to single-token autoregressive decode on CPUs is to co-design the model architecture and the inference runtime together, and presents cflow, a CPU-first streaming engine, alongside a family of pipeline-native transformer architectures whose inter-layer dependency grap...
This work introduces Schur Replay, a scale-selection algorithm that reproduces the GPTQ updates caused by each block scale and scores the resulting block error after accounting for compensation from unquantized columns.
Rui-Ying Ding, Jie Li, Kang He et al.· 0 citations
Vision-language-action (VLA) models have emerged as a key component in embodied AI. Among existing approaches, diffusion-based VLA models achieve superior motion quality and generalization. However, diffusion-based VLA models are compute-intensive and must run at high control frequency, e.g., 50-200 Hz. Thus, it impose...
We describe the architecture, training methodology and inference speedups of Granite 5.0 Turbo CTC, a 470 million parameter encoder-only model with an excellent speed-accuracy tradeoff. The architecture uses pyramidal temporal subsampling within Conformer blocks using strided depthwise convolutions, block-diagonal (chu...
Brian Kingsbury, G. Saon, Masayuki Suzuki et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.