Jul 2026· IEEE International Conference on Cloud Computing· pp. 443-453· 0 citations· 33 references
Abstract
Serverless computing has emerged as an attractive deployment model for deep learning model inference workflows, enabling elastic scaling and fine-grained resource billing across function instances. However, scheduling in this setting introduces a competing-objective challenge: placement decisions simultaneously govern data transfer overhead, determined by whether dependent function instances are co-located, and model loading overhead, determined by whether required model weights are memory-resident on the target node. We present AnchorDL, a joint-cost look-ahead scheduler that minimizes the combined cost of both overheads at each placement decision, with a forward term that avoids greedy suboptimality across adjacent data dependencies. Evaluated against three baselines across chain, fan-in, and fan-out workflow topologies under trace-based workloads, AnchorDL achieves the lowest median end-to-end latency across all evaluated workflows and reduces P90 latency by up to 57.4% against the model-centric baseline in the fan-out workflow. The look-ahead term further contributes substantially beyond greedy joint-cost placement.
Or OrionInfer, an adaptive LLM serving system that aligns inference strategies with real-time demand and introduces three key techniques: runtime switching between data parallelism and tensor parallelism with negligible overhead, an efficient inference pipeline that preserves batching efficiency during parallelism tran...
Jingqi Feng, Guang Yang, Yukai Huang et al.· Proceedings of the 32nd ACM...· 0 citations
CELLServe formalizes SLO-constrained joint resource provisioning as an optimization problem with a dedicated algorithm, and introduces an opportunistic instance merging strategy for decode phase functions to reclaim fragmented resources.
Ze-Jian Wang, Nan Lin, Zi-Nuo Cai et al.· ACM Transactions on Architec...· 0 citations
ServerlessT2I is presented, a serverless-native system that decomposes a T2I workflow into loosely coupled model functions that can be independently managed and scheduled and introduces a fair scheduler for multi-tenant serving.
Xiao-Xiao Jiang, Su-Yi Li, Sheng Yao et al.· arXiv.org· 0 citations
The surge of large language model (LLM) applications on personal devices imposes massive, bursty workloads on cloud serving infrastructure. While prefill-decode disaggregation improves throughput and scalability, memory-bound decode instances often suffer from persistent load imbalance, as output lengths are unknown wh...
Tian-Cheng Zhang, Yulin Chen, Yun-Feng Zhao et al.· 0 citations
ElastiCo is presented, an elastic co-location framework that enables training and inference workloads to safely share GPUs through three integrated mechanisms that decomposes the resulting multi-resource allocation problem into per-job configuration selection subproblems via dynamic per-resource shadow prices.
Jing-Hao Wang, Yi-Hang Zhou, Xiaoyang Sun et al.· 0 citations
This work proposes ScaleSense, a proactive, query-level resource scaling framework that addresses the critical performance-cost trade-off while maintaining low-overhead inference latency, and confirms its practical performance in production deployments.
Yifan Wu, Yuhan Li, Zhenhua Wang et al.· Proceedings of the VLDB Endo...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.