Compute Express Link (CXL) enables cost-effective memory capacity expansion by placing additional tiers behind a coherent fabric. To overcome the performance loss induced by the higher latency of CXL, tiering systems keep hot data in DRAM and demote cold data to the slower tiers. While most research has focused on the...
Sujay Yadalam, Saarth Deshpande, Divyanshu Saxena et al.· Proceedings of the 4th Works...· 0 citations
GPUs are expensive, yet inference-serving GPU clusters remain heavily underutilized. To improve utilization, state-of-the-art systems adopt GPU multiplexing. However, optimizing solely for utilization can counterintuitively increase energy consumption. Designing policies that treat power and energy as first-order metri...
Prasoon Sinha, Dimitrios Liakopoulos, Nathan Lemma et al.· Proceedings of the ACM SIGOP...· 0 citations
It is found that fewer tokens need not mean faster or cheaper execution: on Terminal-Bench with Qwen, policies using roughly one-third as many tokens can take 20-80% longer than the uncompressed agent.
Ritul Satish, Prasoon Sinha, Akiho Kawada et al.· 0 citations
This work presents Oneiros, a dynamic remapping engine for multi-tenant LLM serving that dynamically repurposes GPU memory allocated for model parameters as KV cache capacity, enabling nonblocking, unidirectional parameter transfer.