Sep 2026· Proceedings of the International Conference on Parallel Processing· pp. 381-391· 0 citations· 10 references
Abstract
The performance of single-GPU LLM inference is characterized by strong cross-layer interactions spanning model architecture, runtime scheduling, operator execution, and GPU microarchitecture. Unfortunately, a unified understanding of single-GPU LLM inference bottlenecks is still lacking due to two limitations: the lack of traceable cross-layer attribution and the static characterization of KV-cache usage. To this end, we present a cross-layer analysis approach for single-GPU LLM inference that jointly characterizes latency and memory behavior. However, this requires addressing two challenges: cross-layer attribution and dynamic KV-cache characterization. To address the two challenges, we employ four-layer joint profiling for cross-layer attribution and introduce a logical KV computation metric for dynamic KV-cache characterization, respectively. Finally, we systematically characterize representative dense and Mixture-of-Experts (MoE) models under diverse workloads on a single A100 GPU, identify key bottlenecks across inference phases and architectures, and reveal opportunities for further optimization.
Design rules and a reproducible evaluation protocol are contributed that jointly report quality, memory, and end-to-end speed, and a foundation for automated pipeline search under realistic single-GPU constraints is provided.
SA-Scheduler is presented, a structure-derived bottleneck-aware scheduling framework for multitasking MCM-GPUs that determines chip placement without hardware modification or runtime bottleneck profiling, and provides a principled and scalable foundation for multitasking on future MCM-GPUs.
Tie-Jian Zhang, Guangda Zhang, Lu Wang et al.· ACM Transactions on Design A...· 0 citations
LMTracer is presented, a fine-grained and real-time performance profiling framework for production LLM services that embed profiling logic into the execution through graph-embedded probing and streaming buffered profiling data to CPUs on demand to keep the execution of user kernels uninterrupted.
Wei Liu, Yong-Chao He, Bo-Han Zhao et al.· Proceedings of the ACM SIGOP...· 0 citations
A controlled measurement study of self-hosted LLM inference across edge and near-edge deployment nodes: an NVIDIA Jetson AGX Orin and a near-edge server with CPU-only and GPU-enabled inference modes, highlighting that compute-side inference metrics alone can lead to suboptimal placement for latency-sensitive interactiv...
Maysam Khatib, Moysis Symeonides, Demetris Trihinas et al.· 0 citations
Modern LLM architectures and systems render GPU kernel input shapes increasingly dynamic, e.g., conditional expert routing in MoE and batches with variable-length sequences. This dynamism degrades performance in both expert-tuned kernels and compiler frameworks. We identify the root cause as dynamic cross-SM data depen...
Jing-Kai He, Guang-Da Sun, Tian-Jian Li et al.· Proceedings of the ACM SIGOP...· 0 citations
Weave is presented, to the authors' knowledge the first MoE overlap system that performs fine-grained dynamic SM scheduling - deciding per layer and per GPU by routing results at runtime, and achieves a 2.89x geometric-mean MoE-layer speedup and a 1.33x geometric-mean end-to-end speedup over five state-of-the-art basel...
Ziyu Huang, Yangjie Zhou, Chen-Hao Zhu et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 8, 2026