Cross-Layer Performance Analysis of Single-GPU Large Language Model Inference
This work presents a cross-layer analysis approach for single-GPU LLM inference that jointly characterizes latency and memory behavior, and systematically characterize representative dense and Mixture-of-Experts models under diverse workloads on a single A100 GPU.