Cross-Layer Performance Analysis of Single-GPU Large Language Model Inference
TL;DR
This work presents a cross-layer analysis approach for single-GPU LLM inference that jointly characterizes latency and memory behavior, and systematically characterize representative dense and Mixture-of-Experts models under diverse workloads on a single A100 GPU.
Abstract
The performance of single-GPU LLM inference is characterized by strong cross-layer interactions spanning model architecture, runtime scheduling, operator execution, and GPU microarchitecture. Unfortunately, a unified understanding of single-GPU LLM inference bottlenecks is still lacking due to two limitations: the lack of traceable cross-layer attribution and the static characterization of KV-cache usage. To this end, we present a cross-layer analysis approach for single-GPU LLM inference that jointly characterizes latency and memory behavior. However, this requires addressing two challenges: cross-layer attribution and dynamic KV-cache characterization. To address the two challenges, we employ four-layer joint profiling for cross-layer attribution and introduce a logical KV computation metric for dynamic KV-cache characterization, respectively. Finally, we systematically characterize representative dense and Mixture-of-Experts (MoE) models under diverse workloads on a single A100 GPU, identify key bottlenecks across inference phases and architectures, and reveal opportunities for further optimization.