Skip to content
#large language models Book Open access

Cross-Layer Performance Analysis of Single-GPU Large Language Model Inference

Sep 2026 · Proceedings of the International Conference on Parallel Processing · pp. 381-391 · 0 citations · 10 references

Abstract

The performance of single-GPU LLM inference is characterized by strong cross-layer interactions spanning model architecture, runtime scheduling, operator execution, and GPU microarchitecture. Unfortunately, a unified understanding of single-GPU LLM inference bottlenecks is still lacking due to two limitations: the lack of traceable cross-layer attribution and the static characterization of KV-cache usage. To this end, we present a cross-layer analysis approach for single-GPU LLM inference that jointly characterizes latency and memory behavior. However, this requires addressing two challenges: cross-layer attribution and dynamic KV-cache characterization. To address the two challenges, we employ four-layer joint profiling for cross-layer attribution and introduce a logical KV computation metric for dynamic KV-cache characterization, respectively. Finally, we systematically characterize representative dense and Mixture-of-Experts (MoE) models under diverse workloads on a single A100 GPU, identify key bottlenecks across inference phases and architectures, and reveal opportunities for further optimization.

Read PDF

Similar papers

Open access Aug 2026

Structure-Derived Bottleneck-Aware Scheduling for Multitasking MCM-GPUs

SA-Scheduler is presented, a structure-derived bottleneck-aware scheduling framework for multitasking MCM-GPUs that determines chip placement without hardware modification or runtime bottleneck profiling, and provides a principled and scalable foundation for multitasking on future MCM-GPUs.

Tie-Jian Zhang, Guangda Zhang, Lu Wang et al. · 0 citations
Book Open access Sep 2026

LMTracer: Fine-Grained and Real-Time Performance Profiling for Production LLM Systems

LMTracer is presented, a fine-grained and real-time performance profiling framework for production LLM services that embed profiling logic into the execution through graph-embedded probing and streaming buffered profiling data to CPUs on demand to keep the execution of user kernels uninterrupted.

Wei Liu, Yong-Chao He, Bo-Han Zhao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

A Measurement Study of LLM Inference Trade-offs Across Edge Continuum Hardware

A controlled measurement study of self-hosted LLM inference across edge and near-edge deployment nodes: an NVIDIA Jetson AGX Orin and a near-edge server with CPU-only and GPU-enabled inference modes, highlighting that compute-side inference metrics alone can lead to suboptimal placement for latency-sensitive interactiv...

Maysam Khatib, Moysis Symeonides, Demetris Trihinas et al. · 0 citations
Book Open access Sep 2026

Taming Dynamism on GPUs: Cross-SM Kernel Fusion via SM Cooperation and Just-in-Time Reduction

Modern LLM architectures and systems render GPU kernel input shapes increasingly dynamic, e.g., conditional expert routing in MoE and batches with variable-length sequences. This dynamism degrades performance in both expert-tuned kernels and compiler frameworks. We identify the root cause as dynamic cross-SM data depen...

Jing-Kai He, Guang-Da Sun, Tian-Jian Li et al. · 0 citations
Preprint Sep 2026

Weave: Fine-Grained Dynamic SM Scheduling in an MoE Megakernel for Compute-Communication Overlap

Weave is presented, to the authors' knowledge the first MoE overlap system that performs fine-grained dynamic SM scheduling - deciding per layer and per GPU by routing results at runtime, and achieves a 2.89x geometric-mean MoE-layer speedup and a 1.33x geometric-mean end-to-end speedup over five state-of-the-art basel...

Ziyu Huang, Yangjie Zhou, Chen-Hao Zhu et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.