Sharing GPUs among many deep learning models is crucial for cost-efficient inference, but bursty multi-model workloads can easily overwhelm GPU capacity, causing severe tail latency and SLO goodput drops. Existing solutions—whether traffic-aware scheduling or hardware-level resource partitioning—can only juggle content...
Shi-Jie Peng, Yanying Lin, Cheng-Zhi Lu et al.· Proceedings of the Internati...· 0 citations
Evaluation on a 5-node GPU cluster under synthetic and production-derived churn shows that Connex reduces P99 tail spikes by up to 85% compared to NCCL-based baselines, achieves sub-second cutover, and maintains 100% goodput at moderate loads where baselines collapse to 0–28%, while incurring less than 5% steady-state...
Yanying Lin, Vincent Liu, Tao Luo et al.· Conference on Applications,...· 1 citation
Large language model (LLM) deployment at the network edge faces a fundamental paradox: applications require full-scale models for sophisticated reasoning, yet edge devices impose severe resource constraints across computation, memory, and network. Existing approaches fail to effectively orchestrate resources across the...
Yanying Lin, Baicheng Chen, Xinyu Zhang et al.· International Symposium on C...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.