Sep 2026· Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence· 0 citations· 30 references
TL;DR
This work proposes a Deep Reinforcement Learning-based auto-scaling framework tailored for the PD architecture that enables the agent to capture non-linear load dynamics, thereby achieving decoupled and precise scaling for prefill and decode pools.
Abstract
The disaggregated Prefill-Decode (PD) architecture has emerged as a prominent paradigm for efficient Large Language Model inference serving. However, resource management remains a critical challenge, particularly under the dual burstiness of real-world scenarios—characterized by volatile fluctuations in both request arrival rates and Prompt-to-Response ratios. Existing rule-based heuristics often fail to accurately identify system bottlenecks, leading to severe resource misallocation and Service Level Objective (SLO) violations. To address this, we propose a Deep Reinforcement Learning-based auto-scaling framework tailored for the PD architecture. By modeling the resource allocation problem as a Markov Decision Process, our framework enables the agent to capture non-linear load dynamics, thereby achieving decoupled and precise scaling for prefill and decode pools. Furthermore, to mitigate Head-of-Line blocking caused by scaling latency, we design an immediate rescheduling mechanism that migrates queued tasks to newly ready nodes in real-time. Experimental results driven by Azure bursty load traces demonstrate that our framework significantly reduces computational costs by 25.2% and 28.3% compared to the Static configuration and HeteroScale, respectively, while strictly adhering to SLOs.
CELLServe formalizes SLO-constrained joint resource provisioning as an optimization problem with a dedicated algorithm, and introduces an opportunistic instance merging strategy for decode phase functions to reclaim fragmented resources.
Ze-Jian Wang, Nan Lin, Zi-Nuo Cai et al.· ACM Transactions on Architec...· 0 citations
WAQ-LLM is proposed, a performance optimization framework to find the optimal deployment configuration for multi-instance LLM serving that can provide adaptive and hybrid configurations to handle diverse workloads, surpassing static deployments limited to a single instance type.
Jia-Xin Lai, Yi-Zhou Luo, Qiang Wang· Proceedings of the Internati...· 0 citations
The results indicate that iScavenger provides configurable operating points in the latency–utilization trade-off, limiting additional Sticky-flow RTT while achieving higher background throughput than conservative baseline policies, and highlight the potential of short-term traffic-demand prediction for proactive conten...
Shah M. Emad Uddin, Karl-Johan Grinnemo, Arunselvan Ramaswamy et al.· IEEE Open Journal of the Com...· 0 citations
DOPS (dynamic operator scheduling), a hardware-aware, closed-loop framework that jointly optimizes operator scheduling and blockwise weight layouts and supports systematic analysis of workload sensitivity and hardware scalability for LLM serving is presented.
Jiaqi Yang, Jia-Yi Li, Yihan Fu et al.· arXiv.org· 0 citations
Large-scale rollouts have become a core component of modern LLM systems, spanning reinforcement learning (RL) post-training, on-policy distillation (OPD), and sampling-heavy evaluation pipelines. Unlike online serving, which is typically optimized for request-level latency and throughput, a small number of long-tail ge...
Tianqi Xu, Lu Lv, Hao-Yang Huang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.