Jun 2026
KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding
KernelFlume is presented, a decode-centric architecture that disaggregates the stable projection/FFN path from core-attention computation: weight nodes execute dense projection/FFN kernels, while weightless attention nodes store token-range KV partitions and scale with request-state demand.
Guangyu Xiang, Xueze Kang, Lin Zhang et al.
· arXiv.org · 1 citation