Skip to content
Conference

Performance Characterization of LLM Inference under Limited GPU Resources

Jul 2026 · Annual International Computer Software and Applications Conference · pp. 1018-1023 · 0 citations · 13 references
Computer Science

Abstract

With the widespread adoption of large language models (LLMs), the demand for graphics processing units (GPUs)-essential for accelerating LLM training and inference- has increased significantly. This has led to rising procurement and operation costs, imposing a significant financial burden on research institutions and industry. Efficient utilization of GPU resources has thus become a critical challenge. In this study, we investigate strategies to maximize resource efficiency by hosting multiple models on a single GPU rather than dedicating each GPU to a single model. We examined various GPU resource partitioning approaches to improve the utilization of limited GPU resources. Specifically, we compared two resource allocation methods for concurrently executing two models on a single GPU: using vLLM, a high-performance LLM inference framework, and using NVIDIA Multi-Instance GPU (MIG). The results demonstrated that the MIG configuration increased total throughput by approximately 700-950 tokens/s compared with vLLM-only execution, suggesting that partitioning a GPU into independent MIG instances can improve throughput for concurrent model execution.

View source

Similar papers

Oct 2026

SynergyScale: Optimizing Offloading and Task Partitioning for Efficient Model Training

Deep neural networks (DNNs) with billions of parameters power many important applications, but their training is fundamentally constrained by the limited on-chip memory of GPUs. This memory wall forces training to rely on distributed execution or memory offloading, both of which introduce substantial inefficiencies. Ex...

Xiaoyang Sun, Jie Xu, Zheng Wang · 0 citations
#large language models Book Open access Sep 2026

Cross-Layer Performance Analysis of Single-GPU Large Language Model Inference

The performance of single-GPU LLM inference is characterized by strong cross-layer interactions spanning model architecture, runtime scheduling, operator execution, and GPU microarchitecture. Unfortunately, a unified understanding of single-GPU LLM inference bottlenecks is still lacking due to two limitations: the lack...

Zong-Xing Zhao, Xiaqing Li, Ze-Kai Meng et al. · 0 citations

Performance Evaluation of LLM Inference Engines

This thesis presents a performance evaluation of three widely used open-source inference engines: vLLM, SGLang, and llama.cpp, and summarizes the experimental findings into selection recommendations for practical deployment scenarios, providing a reference for developers and researchers in choosing an appropriate infer...

Unknown authors · 0 citations

K AIROX : Adaptive GPU–CPU Hybrid LLM Inference via Online Neuron Balancing

K AIROX introduces a Live Pipeline designed to prefetch neurons by predicting next-layer activation patterns, a mechanism that dynamically redistributes neurons between the GPU and CPU based on activation patterns, and a Temporal Activation Momentum cache policy to prioritize neurons with sustained utility while minimi...

Yapeng Jiang, Minghao Gan, Zi-Cong Hong et al. · 0 citations
Preprint Aug 2026

LazyTrain: Limited-resource Allocation toward Zero-waste Yield Optimization in Large Language Model Training

LazyTrain is proposed, an optimization layer over a layer-streaming executor that formulates checkpoint selection, activation placement, recomputation, and CPU-GPU-NVMe communication overlap as a mixed-integer scheduling problem, then executes the solved policy during training.

Xiao-Jun Wu, Ce-Hao Yang, Hong-Hao Liu et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.