Caches reduce latency and network traffic, and cache performance largely depends on the eviction policy. Eviction effectiveness is typically measured by byte and object miss ratios. Although learning-based eviction policies can reduce misses, their high computational overhead limits practical deployment. This work pres...
Qian Wang, Chang Xu, Wenbin Zhou et al.· ACM Transactions on Computer...· 0 citations
This work presents a cross-layer analysis approach for single-GPU LLM inference that jointly characterizes latency and memory behavior, and systematically characterize representative dense and Mixture-of-Experts models under diverse workloads on a single A100 GPU.
Zong-Xing Zhao, Xia-Qing Li, Ze-Kai Meng et al.· Proceedings of the Internati...· 0 citations
The performance of single-GPU LLM inference is characterized by strong cross-layer interactions spanning model architecture, runtime scheduling, operator execution, and GPU microarchitecture. Unfortunately, a unified understanding of single-GPU LLM inference bottlenecks is still lacking due to two limitations: the lack...
Zong-Xing Zhao, Xiaqing Li, Ze-Kai Meng et al.· Proceedings of the Internati...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.