Speculative decoding reduces sequential Target model calls by verifying multiple tokens from the Draft model in parallel. Yet KV Cache growth limits long-context serving under constrained GPU memory. Offloading KV to CPU memory relieves this pressure. However, existing offloading schemes restore the full KV history bef...
Fei Li, Song Liu, Shi-Qiang Nie et al.· 0 citations
Shingled Magnetic Recording (SMR) and Interlaced Magnetic Recording (IMR) technologies significantly increase disk storage density by overlapping internal tracks, but the resulting frequent read-modify-write(RMW) operations can cause severe performance jitter. Building Log-Structured Merge Tree (LSM-tree) based key-val...
Fang-Xing Yu, Zhi-Ke Li, Chi Zhang et al.· ACM Transactions on Architec...· 0 citations
In cloud computing, multi-tenant shared storage is widely used for efficiency. ZNS SSDs have become a popular choice due to their high throughput and low latency. However, they still face challenges: the zone structure is not exposed to tenants, making performance control difficult, and the multi-namespace design can l...