Wafer-scale accelerators promise ultra-low-latency AI inference, but current system stacks still carry an unsustainably high cost premium. The reason is that many current and emerging inference techniques, such as batching and MoE, introduce runtime dynamism that existing wafer-scale systems cannot support efficiently....
Cong-Jie He, Le Xu, Zhan Lu et al.· Proceedings of the ACM SIGOP...· 0 citations
Wafer-scale accelerators offer a new scaling point for AI infrastructure, but they also create a new compilation regime: communication cost varies sharply with location, and the space of possible placements and execution schedules is enormous. Existing GPU, distributed, and vendor compilation systems largely retain a s...
Ye-Qi Huang, Cong-Jie He, Hao-Cheng Xiao et al.· Proceedings of the ACM SIGOP...· 0 citations
A remote Spectre attack that leaks a JWT token from a co-located victim worker in the Cloudflare Workers production environment is demonstrated, outperform the existing attack by orders of magnitude.
Martin Schwarzl, Hao-Cheng Xiao, Albert Pedersen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.