A contextual restless multi-armed bandit (CRMAB) framework in which a grid operator requests load reductions without observing internal job-scheduling decisions is proposed, demonstrating the economic potential of data-center flexibility as a grid service and highlighting the importance of high-quality, open-source AI workload traces for developing and evaluating such services.
Abstract
The rapid growth of large-scale AI workloads in data centers has placed increasing pressure on power grids in recent years. Since power systems must continuously balance supply and demand, there is growing interests in leveraging data-center workload flexibility as a grid service. We propose a contextual restless multi-armed bandit (CRMAB) framework in which a grid operator requests load reductions without observing internal job-scheduling decisions. Under index-ability guarantee, each data center or physical machine is modeled as a Markov decision process (MDP) over a cyclic virtual-machine (VM) job queue, with unknown rewards and transition dynamics learned online using Thompson sampling and Whittle-index policies. To improve learning under sparse and noisy observations, the framework augments an adaptive Thompson--Whittle (TW) policy with domain-informed transition priors and gated prior mixing. In baseline experiments, the best adaptive refined variant achieves 91.4\% of the oracle reward after 100 rounds and 96.8\% after 1,000 rounds. Across a 16-setting stress test spanning different state-space sizes and levels of contextual noise, the best refined variant consistently outperforms the original TW policy with high confidence while remaining competitive with EXP4. A graph-based prior further incorporates data-center hardware constraints, including computing-resource limits. Overall, the results demonstrate the economic potential of data-center flexibility as a grid service and highlight the importance of high-quality, open-source AI workload traces for developing and evaluating such services.
A deep reinforcement learning (DRL) framework that jointly co-schedules computing and thermal resources so that a hyperscale data center can operate as a grid-interactive flexible load and supports the evolution of hyperscale data centers from passive electricity consumers toward active, grid-interactive participants i...
This work proposes Drift-Aware Sparse Routing (DRS), a nonstationary sparse contextual routing with multiple knapsack constraints and an optional shadow-audit stream that evaluates a small fraction of prompts on several models.
This work introduces MARS (Monte Carlo Tree Search-based Adaptive and Responsive Scheduler), a training-free HPC scheduler whose optimization goal is configurable through a reward function rather than baked into a learned model.
Yash Kurkure, Yihe Zhang, Zhiling Lan et al.· 0 citations
DeepShare is a scheduler that uses a continuous tenant-assurance signal to coordinate these decisions at runtime to achieve a more advantageous utilization-QoS trade-off than optimizing quotas, scheduling, and resource sharing independently.
Jing-Hao Wang, Yi-Hang Zhou, Xiao Zhou et al.· 0 citations
Cloud computing resource allocation remains a critical challenge, with organizations wasting an estimated $109 billion annually on idle or over-provisioned resources. Traditional allocation strategies—static provisioning, threshold-based autoscaling, and time-series forecasting—fail to capture the complex, non-stationa...
Msr Prasad· International Journal of Tec...· 0 citations
We study a two-server load balancing system with heterogeneous service rates that are a priori unknown to the dispatcher. The goal is to route customers according to the Shortest--Expected--Delay (SED) policy, but this requires knowledge of the service rates. Empirical policies that route based on estimates perform poo...
Sanne van Kempen, Jaron Sanders, Fiona Sloothaak et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.