TurboBus is presented, which pools PCIe bandwidth across co-located jobs via emerging scale-up fabrics and reduces first-token latency by up to 40% for on-demand model loading, achieves up to 1.6x throughput for KV-cache-offloaded inference, and accelerates training by up to 7%, while imposing less than 1% overhead on...
Xinyu Yang, Kaiqiang Xu, Kai Chen· Conference on Applications,...· 0 citations
The proposed Wireless GPU Computing Infrastructure (WiCi) can reduce time to first token by up to 90%, improve the token rate by approximately 39x compared to local inference on mobile devices for the same model, and support much larger models.
Yibin Shen, Wei Li, Kai-Qiang Xu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.