Skip to content
Conference Open access

SplitScaling: Adaptive Scaling for Disaggregated LLM Serving Against Traffic Bursts via DRL

Sep 2026 · Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence · 0 citations · 30 references

TL;DR

This work proposes a Deep Reinforcement Learning-based auto-scaling framework tailored for the PD architecture that enables the agent to capture non-linear load dynamics, thereby achieving decoupled and precise scaling for prefill and decode pools.

Abstract

The disaggregated Prefill-Decode (PD) architecture has emerged as a prominent paradigm for efficient Large Language Model inference serving. However, resource management remains a critical challenge, particularly under the dual burstiness of real-world scenarios—characterized by volatile fluctuations in both request arrival rates and Prompt-to-Response ratios. Existing rule-based heuristics often fail to accurately identify system bottlenecks, leading to severe resource misallocation and Service Level Objective (SLO) violations. To address this, we propose a Deep Reinforcement Learning-based auto-scaling framework tailored for the PD architecture. By modeling the resource allocation problem as a Markov Decision Process, our framework enables the agent to capture non-linear load dynamics, thereby achieving decoupled and precise scaling for prefill and decode pools. Furthermore, to mitigate Head-of-Line blocking caused by scaling latency, we design an immediate rescheduling mechanism that migrates queued tasks to newly ready nodes in real-time. Experimental results driven by Azure bursty load traces demonstrate that our framework significantly reduces computational costs by 25.2% and 28.3% compared to the Static configuration and HeteroScale, respectively, while strictly adhering to SLOs.

Read PDF

Similar papers

Open access Aug 2026

CELLServe: An SLO-Aware and Cost Efficient LLMs Serving System for Serverless Computing Environments

CELLServe formalizes SLO-constrained joint resource provisioning as an optimization problem with a dedicated algorithm, and introduces an opportunistic instance merging strategy for decode phase functions to reclaim fragmented resources.

Ze-Jian Wang, Nan Lin, Zi-Nuo Cai et al. · 0 citations
#large language models Book Open access Sep 2026

WAQ-LLM: Optimizing Multi-Instance LLM Deployment via Workload-Aware Queueing Model

WAQ-LLM is proposed, a performance optimization framework to find the optimal deployment configuration for multi-instance LLM serving that can provide adaptive and hybrid configurations to handle diverse workloads, surpassing static deployments limited to a single instance type.

Jia-Xin Lai, Yi-Zhou Luo, Qiang Wang · 0 citations
Open access 2026

iScavenger: Predictive Multi-Flow Scheduling for Delay-Sensitive Traffic in ATSSS Networks

The results indicate that iScavenger provides configurable operating points in the latency–utilization trade-off, limiting additional Sticky-flow RTT while achieving higher background throughput than conservative baseline policies, and highlight the potential of short-term traffic-demand prediction for proactive conten...

Shah M. Emad Uddin, Karl-Johan Grinnemo, Arunselvan Ramaswamy et al. · 0 citations
Jul 2026

Beyond Prefill-Decode Disaggregation: Dissecting LLM Inference for Heterogeneous Platforms via Dynamic Operator Scheduling

DOPS (dynamic operator scheduling), a hardware-aware, closed-loop framework that jointly optimizes operator scheduling and blockwise weight layouts and supports systematic analysis of workload sensitivity and hardware scalability for LLM serving is presented.

Jiaqi Yang, Jia-Yi Li, Yihan Fu et al. · 0 citations
Preprint Aug 2026

TailSieve: Partial-Rollout-Guided Tail Routing for LLM Rollouts

Large-scale rollouts have become a core component of modern LLM systems, spanning reinforcement learning (RL) post-training, on-policy distillation (OPD), and sampling-heavy evaluation pipelines. Unlike online serving, which is typically optimized for request-level latency and throughput, a small number of long-tail ge...

Tianqi Xu, Lu Lv, Hao-Yang Huang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.