Wafer-scale accelerators promise ultra-low-latency AI inference, but current system stacks still carry an unsustainably high cost premium. The reason is that many current and emerging inference techniques, such as batching and MoE, introduce runtime dynamism that existing wafer-scale systems cannot support efficiently....
Cong-Jie He, Le Xu, Zhan Lu et al.· Proceedings of the ACM SIGOP...· 0 citations
Wafer-scale accelerators offer a new scaling point for AI infrastructure, but they also create a new compilation regime: communication cost varies sharply with location, and the space of possible placements and execution schedules is enormous. Existing GPU, distributed, and vendor compilation systems largely retain a s...
Ye-Qi Huang, Cong-Jie He, Hao-Cheng Xiao et al.· Proceedings of the ACM SIGOP...· 0 citations
This work presents TileSight, a tile-centric performance-modeling tool that elevates the tile from a programming primitive to an analysis primitive, and selects tile configurations competitive with strong vendor and expert baselines.
Zhiwen Mo, Yu Cheng, Lei Wang et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.