Skip to content
Book Open access

Beyond Utilization: Energy-Conscious GPU Sharing for Inference Serving

Sep 2026 · Proceedings of the ACM SIGOPS 32nd Symposium on Operating Systems Principles · 0 citations · 167 references

Abstract

GPUs are expensive, yet inference-serving GPU clusters remain heavily underutilized. To improve utilization, state-of-the-art systems adopt GPU multiplexing. However, optimizing solely for utilization can counterintuitively increase energy consumption. Designing policies that treat power and energy as first-order metrics requires understanding how deployment decisions—GPU allocation size, operating frequency, and batch size—affect energy, latency, and throughput. These relationships are complex, leading existing approaches to rely on extensive profiling. At scale, profiling becomes prohibitively expensive: each model can be deployed under hundreds of configurations, and profiling itself incurs significant energy cost, necessitating accurate yet energy-conscious methods. Further, such systems must model the power draw of colocated models on shared GPUs and adapt to dynamic workload fluctuations. We present EnerTune, an inference serving system that reduces energy consumption while meeting performance SLOs. EnerTune introduces analytical models to estimate per-model performance and power, and the power draw of colocated models on shared GPUs, and uses them in an energy-aware bin-packing algorithm to jointly determine model placement and configuration. EnerTune meets performance SLOs while reducing energy consumption by 1.4-2.3× and power draw by 1.3-2.6× over state-of-the-art baselines.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.