Model commitment (MC) is proposed, a mixed-integer linear programming framework that jointly schedules model deployment and cross-site request routing under power constraints and electricity-price signals and enables AI data center operators to achieve a 100% service rate under time-varying grid conditions and reduce total operating cost.
Abstract
AI data centers may face power supply shortages during certain periods, requiring operators to shift large language model (LLM) inference workloads spatially to maintain service rates. However, existing workload-shifting methods typically assume that any data center with sufficient computing resources can immediately serve shifted requests, which may lead to infeasible transfers and unserved demand. This letter proposes model commitment (MC), a mixed-integer linear programming framework that jointly schedules model deployment and cross-site request routing under power constraints and electricity-price signals. First, MC formulates the intertemporal coupling introduced by model replica loading. Second, it translates prefill and decode latency requirements into the amount of demand that each replica can serve. Case studies based on real-world data show that MC enables AI data center operators to achieve a 100% service rate under time-varying grid conditions and reduce total operating cost by 29.0%.
The rapid growth of artificial intelligence (AI) data centers has introduced new challenges to power system operation. As their power demand becomes larger and more variable, quantitatively characterizing their demand flexibility is increasingly important for effective power system coordination. However, heterogeneous...
RAPID is proposed, a region-aware and power-informed scheduling framework that integrates static and online heuristic schedulers for large-scale AI request scheduling that significantly reduces carbon emissions, electricity costs, and total energy consumption.
B. Ding, Cai-Ning Wang, Ka-Fei Tang et al.· Sustainability· 0 citations
PowerSlider does so with a new Flex SLO contract that turns bounded user slack into an optimization constraint, prefill--think--answer disaggregation exposing per-stage frequency and KV control, and a Karush--Kuhn--Tucker (KKT) online solver re-solving within 7.7 ms of every cap change.
Yueying Li, Jia-Yang Chen, Yuan-Fan Chen et al.· 2 citations
Data centers can provide demand-side flexibility by shifting adjustable computing workloads, but price-aware scheduling must preserve service completion and control delay loss. This paper proposes reliability-oriented electricity-computing coordination for data centers. Under a seven-day minute-level setting, an FCFS-l...
Hong-Da Gao, Bao-Di Han· 2026 8th International Confe...· 0 citations
Recent advances in large language models (LLMs) are driving the emergence of multi-modal and agentic services for mobile users through cloud and edge infrastructures, where long-context workloads pose daunting challenges for inference latency. Existing disaggregated LLM serving systems largely rely on hardware profilin...
Shi-Cong Liu, Xiang-Hao Yu, Zheng-Run Gao et al.· 0 citations
Securing grid interconnection capacity has become a bottleneck for AI data center projects and can take longer than constructing the facilities themselves. This mismatch can delay deployment for years, making early interconnection planning essential. This paper develops ICP-AI, an interconnection capacity planning fram...
Hassan Zahid Butt, Rida Fatima, Xing-Peng Li· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.