TeDiServe: High SLO Attainment Serving for Diffusion Language Models
TeDiServe enables deadline-aware scheduling and adaptive load control through confidence-threshold adjustment, and dynamically reconfigures the cluster by solving a quality-aware optimization problem, while explicitly modeling the step-level heterogeneity introduced by approximate KV caching.
Tzu-Tao Chang, Benjamin Yuanyang Hong, Kiet C. Pham et al.
· 1 citation