Omni models unify text, speech, image, and multimodal reasoning in a single serving backend, but this unified deployment exposes a new scheduling problem. Requests with different output modalities may share an initial multimodal backbone and then diverge into downstream generation stages, creating heterogeneous first-response metrics and service-level objective (SLO) targets on the same GPU. Existing large language model (LLM) and multimodal serving systems mainly optimize token progress or input-side processing, and they do not jointly control temporal sharing in the shared stage and spatial sharing among co-running stages. This paper presents HorizonServe, a single-GPU omni-model serving system that coordinates request admission and GPU allocation under heterogeneous SLOs. HorizonServe profiles per-class first-response latency, protects requests with limited slack, rotates shared-stage opportunities across execution paths, and throttles the shared-stage streaming multiprocessor (SM) allocation when downstream stages are active. Across three omni-model workloads and two GPU platforms, HorizonServe improves SLO attainment by up to 4.9$\times$ in arrival-rate sweeps and 7.0$\times$ under downstream-heavy traffic, and reduces per-class first-response latency by 38.4--63.7\%.
Multi-model LLM serving is moving toward shared MaaS clusters, where co-hosted models compete for a fixed GPU budget while each model experiences time-varying demand and must satisfy its own latency SLO. Existing LLM autoscalers remain largely model-local: their signals expose local runtime activity or delayed latency...
Xin Zhang, Xian-Yan Xie, Zhen He et al.· 0 citations
In this paper, we study a mixed-prompt scenario—where both short and long prompts coexist—in an LLM inference serving system that supports diverse applications with heterogeneous iteration-time SLOs. To improve throughput for long prompts, prior work divides them into chunks and batches requests or chunks to meet the t...
Hai-Ying Shen, Tanmoy Sen, Yuxiong He· Proceedings of the Internati...· 0 citations
Large Language Models (LLMs) are increasingly deployed in latency-sensitive applications, where real-time serving must satisfy stringent service-level objectives (SLOs). However, request intensities fluctuate over time, and under low load LLM services leave a substantial fraction of GPU task idle. A promising approach...
Online image generation with Diffusion Transformers (DiTs) must meet latency service-level objectives (SLOs) while using GPU resources efficiently. Existing systems improve GPU utilization by batching multiple requests for joint execution. However, request-level batching offers limited control over batch size: batches...
Zhexiang Zhang, Min-Chen Yu, Yi-Fan Sun et al.· 0 citations
In this paper, we present a latency-aware scheduler for large-language-model (LLM) inference across mobile devices, edge servers, and a remote cloud. Our fine-grained delay model captures OFDMA uplink/downlink rates, KV-cache backhaul serialization, and profiled GPU planning-chunk resource constraints, enabling per-req...
Xinghan Wang, Xiao-Xiong Zhong, Wei-Hong Yang et al.· IEEE Transactions on Paralle...· 0 citations
PackServe is a scheduler designed to reduce resource costs while meeting latency SLOs for agentic LLM serving, which uses compact white-box models to predict latency under prefill/decode interference and packs requests onto fewer serving instances while preserving KVC reuse and SLO constraints.
Zhi-Yuan Tan, De-Jiang Zhu, Jing-Zhe Jiang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.