Preprint
Jul 2026
A CXL Memory Rack for Multi-Turn LLM Serving
This paper presents HyMCache, a CXL memory rack for multi-turn LLM serving using cost-efficient CXL-hybrid memory, which combines a small amount of in-device DRAM with large SSD-backed capacity behind a CXL interface to efficiently support TB-scale SSD-backed KV reuse.
Hakbeom Jang, Inho Song, Hoshik Kim et al.
· 0 citations