Skip to content
Preprint

CLASP: Chained-Request-Aware Scaling and Operator Placement for Serverless Stream Processing

Aug 2026 · 0 citations · 44 references
Computer Science

TL;DR

CLASP, a scaling and scheduling strategy for stream processing in stateful serverless environments, which improves throughput by up to 3.3x and reduces median end-to-end latency by up to 76% compared with state-of-the-art scaling strategies.

Abstract

Stateful serverless (Function-as-a-Service) environments, whose workers host state servers, are increasingly used for stream processing. A stream application is a pipeline of operators, where each operator forwards intermediate data downstream through a chained request. As input rates fluctuate, the system should adjust operator parallelism and place instances across workers to sustain the incoming rate. Existing approaches do so without fully accounting for chained-request overhead, leading them to misestimate the required number of workers. Too few leave the cluster unable to keep up with the input rate, while too many route a larger fraction of chained requests across worker boundaries, increasing end-to-end latency. We propose CLASP, a scaling and scheduling strategy for stream processing in stateful serverless environments. At runtime, CLASP estimates execution cost and chained-request cost from observed metrics. Under a capacity model that covers the two costs, it adjusts operator parallelism and packs operators onto the fewest workers that can sustain the target input rate. Once a scaling decision is made, CLASP migrates each operator's state together with its instances, thereby minimizing execution pause time. Experiments show that CLASP improves throughput by up to 3.3x and reduces median end-to-end latency by up to 76% compared with state-of-the-art scaling strategies.

View source

Similar papers

Preprint Sep 2026

Adaptive Context Parallelism for Production LLM Serving

As LLM context windows expand and input sequences grow longer, serving systems face increasing computational and memory demands. Context parallelism (CP), which partitions the input sequence across multiple ranks to parallelize the computation, has therefore become increasingly important for efficient LLM serving. Howe...

Jiarui Guo, Rong-Le Wang, Pei-Jun Huang et al. · 0 citations
Book Open access Sep 2026

FaaSHive: Addressing Limited Visibility in Function-as-a-Service via Worker-driven Scheduling

Effective request placement in Function-as-a-Service (FaaS) platforms requires timely visibility into worker state, which changes rapidly as workers create, reuse, and evict short-lived function instances. Under the conventional cloud manager–worker architecture commonly used in FaaS platforms, however, this visibility...

Seonggyu Han, Sangwoo Kim, Minho Kim et al. · 0 citations
Preprint Aug 2026

Epico: Long-Lived WebAssembly Components for High-Performance Serverless Stream Processing

While serverless computing is popular, its dominant Function-as-a-Service (FaaS) model is ill-suited for stream processing because its stateless, centrally orchestrated functions cannot efficiently handle continuous, low-latency event flows. We introduce Epico, a serverless runtime explicitly designed to resolve these...

Matteo Della Bartola, Valerio Besozzi, Patrizio Dazzi et al. · 0 citations
Open access Aug 2026

CELLServe: An SLO-Aware and Cost Efficient LLMs Serving System for Serverless Computing Environments

CELLServe formalizes SLO-constrained joint resource provisioning as an optimization problem with a dedicated algorithm, and introduces an opportunistic instance merging strategy for decode phase functions to reclaim fragmented resources.

Ze-Jian Wang, Nan Lin, Zi-Nuo Cai et al. · 0 citations
Preprint Aug 2026

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving

OpScale is presented, a practical operator-level orchestration framework of profiling, provisioning, placement, and runtime serving that attains SLOs with up to 36.3% fewer GPUs and 28% less power, or achieves 44% higher throughput under fixed cost budgets.

Xingqi Cui, Chieh-Jan Mike Liang, Ziang T. Tang et al. · 0 citations
Book Open access Sep 2026

Janus: Multi-LLM Serving at Production Scale

Production LLM serving multiplexes hundreds of heterogeneous models on shared clusters, exposing three challenges that existing systems fail to address simultaneously: unpredictable bursts, power-law application popularity, and heterogeneous yet complementary resource demands. We present Janus, a Service-Engine co-desi...

Tian-Bao Zhou, Yi Wang, Yu Zhou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.