AI and HPC infrastructure increasingly serves workload portfolios that combine dense tensor computation, sparse kernels, large memory footprints, and communication-intensive collectives. Supporting these portfolios requires coordinated choices across accelerators, memory tiers, scale-up fabrics, and cluster networks. T...
Yuchen Fan, Minghong Sun, Jikui Ma et al.· arXiv.org· 0 citations
H3-Attn is proposed, an Attention-efficient 3D DRAM PNM processor for low-batch LLM inference that features a hybrid head parallelism for Attention processing, whereby various optimized Attention mechanisms with spatial tiled FlashAttention can be flexibly enabled with fully leveraged 3D DRAM PNM bandwidth.
Yao-Lei Li, Wen-Bin Jia, Zhan-Chen Zhao et al.· International Symposium on L...· 0 citations
Diffusion models have shown marked advancements in 2-D generation and have also become focal points in 3-D generation via consistent multiview image generation. However, the computation and memory demands hinder their real-time deployment on mobile and edge devices. Moreover, the reduction of diffusion timesteps leads...
Wenxun Wang, Li-Kai Ma, Chen Tang et al.· IEEE Transactions on Compute...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.