DASched: Dependency-Aware Scheduling for Multi-Stage Pipelines under Shared GPU Resources
Abstract
Modern AI applications increasingly rely on multi-stage pipelines that link multiple computational stages into end-to-end workflows, such as video analytics, recommendation systems, and healthcare analysis. Existing serving systems optimize each stage in isolation, overlooking execution-level dependencies that arise when stages share GPU resources. Imbalanced batching and GPU partitioning across interdependent stages cause execution stalls, idling GPU resources and amplifying end-to-end latency. This paper presents DASched, a dependency-aware scheduling framework that jointly coordinates temporal batching, spatial GPU partitioning, and structural morphing for balanced multi-stage execution. It integrates three mechanisms: dynamic dependency-aware batching for temporal alignment, adaptive GPU allocation for spatial balance, and elastic pipeline morphing for scalable resource efficiency. Experiments on diverse multi-stage pipelines show that DASched improves throughput by up to 1.76 × , while reducing GPU usage and latency by up to 71% and 50%, respectively, demonstrating the effectiveness of dependency-aware scheduling for shared-GPU inference.