Skip to content
Preprint

Beyond Scalar Flexibility: From Eligible AI Workloads to Dependable Load Relief

Sep 2026 · 0 citations · 33 references
Engineering Computer Science

TL;DR

This work reconstructs 4,439 hourly power observations from a 185-day trace of 155,410 GPUs and derives a workload-semantic flexibility envelope that can be written into interconnection and demand-response contracts.

Abstract

Grid studies often represent data-center flexibility as a fixed percentage of load, although no public production trace has shown how much eligible load persists across event durations or co-moves across clusters. We reconstruct 4,439 hourly power observations from a 185-day trace of 155,410 GPUs and derive a workload-semantic flexibility envelope. The fleet's time-averaged Monte Carlo median facility demand is 55.8 MW, while immediate eligible curtailment averages 3.55 MW after retaining allocated-GPU idle power: 12.1% of workload power and 6.35% of median facility power. Under full realization of that eligibility, 95%-available relief falls from 2.51 MW for one hour to 2.32 MW for four hours and 1.95 MW for 24 hours; a common realizable fraction q scales every value exactly by q. A mean-calibrated scalar overstates these quantities by 17%, 25%, and 47%, while a scalar tail-calibrated at four hours understates the one-hour product by 6% and overstates the 24-hour product by 17%; the share that reproduces the surface varies by a factor of 1.6 across durations and reliability levels. Aggregating 13 clusters raises four-hour firmness from 0.38 to 0.66, but cross-cluster covariance limits the gain. The production scheduler exposes almost no additional delay-based capacity: newly deferrable arrivals average 0.008 MW and have zero 95%-available capacity. These results replace an assumed flexibility percentage with duration, reliability, portfolio, and realizability terms that can be written into interconnection and demand-response contracts.

View source

Similar papers

Preprint Oct 2026

Ofan: Optimal Load Balancing for AI Training

The extreme collective completion time (CCT) demands of AI workloads challenge existing packet spraying algorithms, which can have trouble efficiently load-balancing workloads that are sent at full line rates. We trace this to a structural cause: on a fat tree, once a packet picks its upward path, the downward path to...

Sarah McClure, E. Cohen, Jakob Krebs et al. · 0 citations
Preprint Sep 2026

ContinuumBench: Benchmarking Joint Autoscaling and Placement Across Evaluation Regimes in the Cloud-Edge Continuum

ContinuumBench is presented, a benchmark that controls workload, connectivity, and calibration assumptions and metrics over completed tasks hide unfinished work in cloud-edge controllers and compares placement-only and scale-capable controllers under declared regimes and stressors.

Lan-Pei Li, Antonino Vaccarella, Vincenzo Lomonaco et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Characterizing Job Power Elasticity for Power-Flexible AI Training

The first systematic characterization of job power elasticity (the sensitivity of throughput to power reductions) in LLM training is presented, and the Power Flexibility Index (PFI) is introduced, a normalized metric that quantifies the performance cost of power reductions and provides a control primitive for SLA-aware...

Philip Colangelo, Charles Dawson, Shayan Sengupta et al. · 0 citations
Book Open access Sep 2026

Taming Inference Workloads at Global Scale: Foundation Model Serving in Amazon Bedrock

Foundation model (FM) inference platforms must manage scarce accelerator capacity distributed unevenly across regions. They must also handle requests whose token consumption varies widely and may be revealed progressively during generation, while sharing capacity across workloads with different latency and throughput o...

Pratik Pankaj Raichura, Somu Perianayagam, Rama Krishna Sandeep Pokkunuri et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.