DASched: Dependency-Aware Scheduling for Multi-Stage Pipelines under Shared GPU Resources
Modern AI applications increasingly rely on multi-stage pipelines that link multiple computational stages into end-to-end workflows, such as video analytics, recommendation systems, and healthcare analysis. Existing serving systems optimize each stage in isolation, overlooking execution-level dependencies that arise wh...