Skip to content
Book Open access

Cross-Layer Performance Analysis of Single-GPU Large Language Model Inference

Sep 2026 · Proceedings of the International Conference on Parallel Processing · 0 citations · 41 references

TL;DR

This work presents a cross-layer analysis approach for single-GPU LLM inference that jointly characterizes latency and memory behavior, and systematically characterize representative dense and Mixture-of-Experts models under diverse workloads on a single A100 GPU.

Abstract

The performance of single-GPU LLM inference is characterized by strong cross-layer interactions spanning model architecture, runtime scheduling, operator execution, and GPU microarchitecture. Unfortunately, a unified understanding of single-GPU LLM inference bottlenecks is still lacking due to two limitations: the lack of traceable cross-layer attribution and the static characterization of KV-cache usage. To this end, we present a cross-layer analysis approach for single-GPU LLM inference that jointly characterizes latency and memory behavior. However, this requires addressing two challenges: cross-layer attribution and dynamic KV-cache characterization. To address the two challenges, we employ four-layer joint profiling for cross-layer attribution and introduce a logical KV computation metric for dynamic KV-cache characterization, respectively. Finally, we systematically characterize representative dense and Mixture-of-Experts (MoE) models under diverse workloads on a single A100 GPU, identify key bottlenecks across inference phases and architectures, and reveal opportunities for further optimization.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.