Pipeline-Native Transformers: Co-Designing Model Architecture and CPU Inference for Bandwidth-Efficient Autoregressive Decode
This report argues that the most effective response to single-token autoregressive decode on CPUs is to co-design the model architecture and the inference runtime together, and presents cflow, a CPU-first streaming engine, alongside a family of pipeline-native transformer architectures whose inter-layer dependency grap...