Skip to content
Conference

AnchorDL: Dual-Locality-based Scheduling for Serverless Inference Workflows

Jul 2026 · IEEE International Conference on Cloud Computing · pp. 443-453 · 0 citations · 33 references

Abstract

Serverless computing has emerged as an attractive deployment model for deep learning model inference workflows, enabling elastic scaling and fine-grained resource billing across function instances. However, scheduling in this setting introduces a competing-objective challenge: placement decisions simultaneously govern data transfer overhead, determined by whether dependent function instances are co-located, and model loading overhead, determined by whether required model weights are memory-resident on the target node. We present AnchorDL, a joint-cost look-ahead scheduler that minimizes the combined cost of both overheads at each placement decision, with a forward term that avoids greedy suboptimality across adjacent data dependencies. Evaluated against three baselines across chain, fan-in, and fan-out workflow topologies under trace-based workloads, AnchorDL achieves the lowest median end-to-end latency across all evaluated workflows and reduces P90 latency by up to 57.4% against the model-centric baseline in the fan-out workflow. The look-ahead term further contributes substantially beyond greedy joint-cost placement.

View source

Similar papers

Book Open access Aug 2026

OrionInfer: Low-Overhead Parallelism Switching and Live Migration for Efficient LLM Serving

Or OrionInfer, an adaptive LLM serving system that aligns inference strategies with real-time demand and introduces three key techniques: runtime switching between data parallelism and tensor parallelism with negligible overhead, an efficient inference pipeline that preserves batching efficiency during parallelism tran...

Jingqi Feng, Guang Yang, Yukai Huang et al. · 0 citations
Open access Aug 2026

CELLServe: An SLO-Aware and Cost Efficient LLMs Serving System for Serverless Computing Environments

CELLServe formalizes SLO-constrained joint resource provisioning as an optimization problem with a dedicated algorithm, and introduces an opportunistic instance merging strategy for decode phase functions to reclaim fragmented resources.

Ze-Jian Wang, Nan Lin, Zi-Nuo Cai et al. · 0 citations
Jul 2026

ServerlessT2I: Efficient Text-to-Image Workflow Serving on a Serverless Platform

ServerlessT2I is presented, a serverless-native system that decomposes a T2I workflow into loosely coupled model functions that can be independently managed and scheduled and introduces a fair scheduler for multi-tenant serving.

Xiao-Xiao Jiang, Su-Yi Li, Sheng Yao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MAPS: Memory-Aware Predictive Scheduling Framework for Large Language Model Serving

The surge of large language model (LLM) applications on personal devices imposes massive, bursty workloads on cloud serving infrastructure. While prefill-decode disaggregation improves throughput and scalability, memory-bound decode instances often suffer from persistent load imbalance, as output lengths are unknown wh...

Tian-Cheng Zhang, Yulin Chen, Yun-Feng Zhao et al. · 0 citations
Preprint Aug 2026

ElastiCo: Elastic Configuration and Interference-Aware Orchestration for GPU Clusters

ElastiCo is presented, an elastic co-location framework that enables training and inference workloads to safely share GPUs through three integrated mechanisms that decomposes the resulting multi-resource allocation problem into per-job configuration selection subproblems via dynamic per-resource shadow prices.

Jing-Hao Wang, Yi-Hang Zhou, Xiaoyang Sun et al. · 0 citations
Open access Aug 2026

ScaleSense: Cost-Intelligent Scaling Framework via Learned Resource Estimation in Alibaba AnalyticDB

This work proposes ScaleSense, a proactive, query-level resource scaling framework that addresses the critical performance-cost trade-off while maintaining low-overhead inference latency, and confirms its practical performance in production deployments.

Yifan Wu, Yuhan Li, Zhenhua Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.