Skip to content
Preprint

PowerSlider: Exploiting Phase Asymmetry for LLM Serving under Demand Response

Aug 2026 · 2 citations · 84 references
Computer Science

TL;DR

PowerSlider does so with a new Flex SLO contract that turns bounded user slack into an optimization constraint, prefill--think--answer disaggregation exposing per-stage frequency and KV control, and a Karush--Kuhn--Tucker (KKT) online solver re-solving within 7.7 ms of every cap change.

Abstract

AI inference clusters are increasingly constrained by instantaneous power, not just energy: grid operators condition new capacity on demand response, imposing time-varying power caps. Existing LLM serving systems optimize a static energy objective or shed fixed priority tiers under load; either way, goodput collapses when the power envelope moves. An LLM pipeline is not a uniform load: compute-bound prefill loses throughput almost linearly with GPU frequency, memory-bound answer decode sustains it down to $0.57\times$ nominal, and reasoning's thinking phase couples KV-cache capacity to scheduling -- so a cap should be steered to where each watt costs the least performance. PowerSlider does so with a new Flex SLO contract that turns bounded user slack into an optimization constraint, prefill--think--answer disaggregation exposing per-stage frequency and KV control, and a Karush--Kuhn--Tucker (KKT) online solver re-solving within 7.7 ms of every cap change, backed by a consolidated fail-safe that power-gates drained instances when DVFS bottoms out on static power. On SGLang with production traces, \sys{} sustains 78.3\% online goodput at a 30\% cap reduction versus 47.6\% for the best of five baselines ($1.64\times$), holds latency-critical tails within $1.3\times$ of nominal (baselines: $2.3$--$6\times$, up to $12\times$), and delivers 92\% mean goodput through a replayed CAISO grid-emergency day bottoming at $0.41\times$ (54\% at the trough; every baseline below 7\%).

View source

Similar papers

Jul 2026

PowerAtlas: Towards Electricity-Computing Co-Scheduling for Power Systems

The rapid growth of AI workloads is turning data centers into large-scale, volatile, yet spatiotemporally flexible grid loads, creating an urgent need for coordinated electricity-computing scheduling. Under stringent grid constraints, schedules from general-purpose large language models (LLMs) are often infeasible, cau...

Kaiwen Jiang, Si-Ya Xu, Zi-Yue Zhu et al. · 0 citations
Preprint Sep 2026

Spatial LLM Workload Shifting Needs Foresight: Model Commitment for AI Data Center Operation under Power Grid Constraints

Model commitment (MC) is proposed, a mixed-integer linear programming framework that jointly schedules model deployment and cross-site request routing under power constraints and electricity-price signals and enables AI data center operators to achieve a 100% service rate under time-varying grid conditions and reduce t...

Bo-Jun Du, Hong-Yang Jia, Tong-Hui Li et al. · 0 citations
Preprint Aug 2026

Smoothing the Ramp, Not the Peak: Scheduling-Induced Power Dynamics of LLM Inference and Their Grid-Scale Consequences

Large language model (LLM) inference serving is a fast-growing electricity load whose power dynamics remain uncharacterized from a grid-planning perspective. Using real, measured GPU power traces, we show that chunked prefill scheduling, a latency-motivated technique already deployed by default in production LLM servin...

Pan-Sheng Li, Yize Chen, Xia Miao et al. · 0 citations
#machine learning Preprint Sep 2026

Phase-Decoupled, Model-Calibrated Power Control for Disaggregated LLM Serving

Datacenter GPU power is the binding constraint on LLM serving capacity, and production serving has shifted to prefill/decode (PD) disaggregation. Deploying NVIDIA's Max-Q inference profile on a disaggregated B200 system, we found its realized gain modest (+8.6% tokens/J), model-dependent, and carrying a mean end-to-end...

Jae Gon Kim, Donghoon Yoo, Hanyul Ryu et al. · 0 citations
Preprint Aug 2026

Routing LLM Inference to the Cleanest Grid in Real Time

Large-language-model inference is a fast-growing electricity load whose marginal carbon intensity varies by more than an order of magnitude across grid regions and across the day, making request placement an attractive lever: no retraining, no hardware change. We report a live validation of carbon-aware inference routi...

Aleks Bernhard, Arif Baran Yardimci · 1 citation
Preprint Sep 2026

Multi-Scale Datacenter Power Modulation

Cloud datacenters must increasingly modulate power in response to time-varying grid and infrastructure constraints. We study this problem as finite-horizon control of a networked hybrid dynamical system, where datacenter power and service capacity depend on interactions between servers, workers, and hosted services. Po...

Akshay Sreekumar, Nicolas H. Christianson, Fiodar Kazhamiaka et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.