Jul 2026· IEEE International Conference on Cloud Computing· pp. 268-278· 0 citations· 27 references
Abstract
Large Language Model (LLM) inference services increasingly rely on heterogeneous GPU clusters to balance cost and performance. However, request routing in such environments is challenging because schedulers must account for hardware heterogeneity, dynamic workload characteristics, and bursty arrivals. Existing approaches either ignore hardware differences, rely on static heterogeneity-aware allocation, or use queue-based proxies that fail to capture the true remaining work under load. We present BYSTANDER, a prediction-based scheduling framework that uses a small language model (SLM) to estimate End-to-End (E2E) latency of request for each GPU pool from request features and current pool state. BYSTANDER then adaptively narrows the candidate pools via Fisher–Jenks grouping and performs queue-aware selection within that set. Its pool-based design keeps prediction overhead scalable as cluster size grows. We evaluate BYSTANDER on ShareGPT and LMSYS-Chat workloads over heterogeneous clusters of 3–7 GPUs (RTX 3090/4090/5090). Compared with round-robin, oracle-weighted round-robin, shortest-queue-first, and SLM Adaptive baselines, BYSTANDER reduces P99 E2E latency by up to 63.4% and P99 time-to-first-token by up to 87.4% under bursty traffic.
Load balancers in practice often rely on fixed heuristics such as weighted round-robin (WRR) or least connection (LC). Although these methods scale well, they do not capture differences in backend service capacity or runtime performance variations. which can increase tail latency and request drop rates in shared cluste...
Hai Pham Thanh, Dang Hoang Nguyen, Anh Nguyen Tuan et al.· IEEE International Conferenc...· 0 citations
The widespread adoption and strong generalizability of large language models (LLMs) lead to highly heterogeneous workloads that exhibit substantial variability in request lengths and latency requirements. This pronounced heterogeneity causes existing scheduling strategies to suffer from head-of-line blocking and ineffi...
Tian-Nan Fu, Jianxiong Liao, Xu Chen et al.· Fall Joint Computer Conferen...· 0 citations
Cascade, an LLM serving system that estimates and continuously updates this per-request latency budget from request characteristics, KV-cache state, and current system load, and uses a single per-request budget to jointly coordinate request scheduling and KV-cache management across the memory hierarchy.
Muhammad Adnan, R. Mahapatra, Prashant J. Nair et al.· 0 citations
CELLServe formalizes SLO-constrained joint resource provisioning as an optimization problem with a dedicated algorithm, and introduces an opportunistic instance merging strategy for decode phase functions to reclaim fragmented resources.
Zejian Wang, Nan Lin, Zinuo Cai et al.· ACM Transactions on Architec...· 0 citations
Or OrionInfer, an adaptive LLM serving system that aligns inference strategies with real-time demand and introduces three key techniques: runtime switching between data parallelism and tensor parallelism with negligible overhead, an efficient inference pipeline that preserves batching efficiency during parallelism tran...
Jingqi Feng, Guang Yang, Yukai Huang et al.· Proceedings of the 32nd ACM...· 0 citations
DeltaServe is presented, a host-agnostic co-serving design that converts this idle inference capacity into LoRA fine-tuning throughput while preserving inference service-level objectives (SLOs).
Jiaxuan Chen, Jianshu She, Ye Yuan et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.