Skip to content
Preprint

Pipeline-Native Transformers: Co-Designing Model Architecture and CPU Inference for Bandwidth-Efficient Autoregressive Decode

Aug 2026 · 0 citations
Computer Science

TL;DR

This report argues that the most effective response to single-token autoregressive decode on CPUs is to co-design the model architecture and the inference runtime together, and presents cflow, a CPU-first streaming engine, alongside a family of pipeline-native transformer architectures whose inter-layer dependency graphs are constructed to permit a vertical, stage-major execution schedule.

Abstract

Single-token autoregressive decode on CPUs is bound by memory bandwidth, not arithmetic: a modern CPU sustains roughly 1 TFLOP/s of compute but only about 50 GB/s from main memory, and each generated token must stream every active weight once. This report argues that the most effective response is to co-design the model architecture and the inference runtime together. It presents cflow, a CPU-first streaming engine, alongside a family of pipeline-native transformer architectures whose inter-layer dependency graphs are constructed to permit a vertical, stage-major execution schedule. cflow stores weights as L2-sized tiles in compute-consumption order, reads only the top-k experts of each mixture-of-experts layer, fuses projections, and executes a delay-aware schedule from per-model dependency parameters. Across five architectures trained on TinyStories, one (arch2_4_combined) achieves a 2.00x reduction in critical-path weight bandwidth (9.00 to 4.50 MB/token) within 0.24 perplexity of the best candidate, and the tile layout incurs 7.29x fewer L1-data read misses than a row-major baseline. On a 30.9-billion-parameter pipeline-native MoE, cflow decodes at 5.94 tokens/s (tok/s) on a 32-vCPU Ice Lake server, ahead of llama.cpp (4.75) and the vLLM CPU backend (1.65) on comparably sized dense models. Realizing the expert-delay window as asynchronous I/O overlap on a disk-resident expert tier yields a further net win of up to 1.68x, matching the overlap model within 1%. Measurement refutes one of the eight design claims and leaves a second inconclusive; both are reported in full, with the conditions under which they would hold.

View source

Similar papers

Preprint Sep 2026

The World Model Hardware Accelerator

Diffusion transformers invert the arithmetic that autoregressive decoding made familiar. There is no token-by-token recurrence: every denoising step is a full-sequence forward pass over static shapes, so the entire schedule is known at compile time and the only serial dimension is the step count itself. We exploit that...

Shashank Chaurasia · 0 citations
#graph neural networks Open access Sep 2026

PipeGNN: A Bandwidth-Efficient GNN Accelerator with Node-Level Pipelined Push Execution

Graph Neural Networks (GNNs) have become a fundamental tool for learning over graph-structured data. Under the message-passing framework, mainstream GNN models alternate between feature transformation and neighborhood aggregation. Fusing these two phases into a node-level pipelined push dataflow, in which each node’s t...

Shi Chen, Jun-Sheng Chang, Yang Guo et al. · 0 citations
#graph neural networks Book Open access Sep 2026

Memory-Aware Joint Optimisation of Partitioning and Scheduling for Pipeline-Parallel Training

OptPipe is presented, a unified framework that jointly optimises partitioning and scheduling for pipeline parallelism and introduces a memory-aware directed acyclic graph (DAG) that captures both task dependencies and the lifetime of intermediate tensors, enabling explicit reasoning about the trade-off between executio...

Ning Wang, A. Raith, Oliver Sinnen · 0 citations
Preprint Aug 2026

Every Expert Counts: ExactMoE for Memory-Efficient W4A16 Inference

ExactMoE, an inference design that applies symmetric group-128 four-bit weight quantization only to routed experts, stores those experts in kernel-native MARLIN form in pinned host memory, and executes all selected experts through a configurable GPU-resident slot cache and fused grouped MoE kernels, identifies a practi...

Amjad Saab · 0 citations
Preprint Aug 2026

MonoMoE: An Efficient Fused Mega-kernel for Quantized MoE Decoding

Mixture-of-Experts (MoE) layers increase model capacity without proportionally increasing arithmetic, but their sparse expert computation is difficult to execute efficiently during autoregressive decode. Existing grouped and batched GEMMs are token-major: they construct expert-local token tiles and obtain parallelism f...

Yu Gong, Kailash Budhathoki, Taeho Kim et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.