Skip to content
Book Open access

DASched: Dependency-Aware Scheduling for Multi-Stage Pipelines under Shared GPU Resources

Sep 2026 · Workshop Proceedings of the 55th International Conference on Parallel Processing · 0 citations · 2 references

Abstract

Modern AI applications increasingly rely on multi-stage pipelines that link multiple computational stages into end-to-end workflows, such as video analytics, recommendation systems, and healthcare analysis. Existing serving systems optimize each stage in isolation, overlooking execution-level dependencies that arise when stages share GPU resources. Imbalanced batching and GPU partitioning across interdependent stages cause execution stalls, idling GPU resources and amplifying end-to-end latency. This paper presents DASched, a dependency-aware scheduling framework that jointly coordinates temporal batching, spatial GPU partitioning, and structural morphing for balanced multi-stage execution. It integrates three mechanisms: dynamic dependency-aware batching for temporal alignment, adaptive GPU allocation for spatial balance, and elastic pipeline morphing for scalable resource efficiency. Experiments on diverse multi-stage pipelines show that DASched improves throughput by up to 1.76 × , while reducing GPU usage and latency by up to 71% and 50%, respectively, demonstrating the effectiveness of dependency-aware scheduling for shared-GPU inference.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.