SlideDP is presented, a synchronous data-parallel runtime for shared-host multi-GPU systems that maintains one authoritative host state, decouples communication routes from state layout, and pipelines parameter delivery, gradient aggregation, and CPU updates across ranks and chunks.
Rui-Jia Yang, Shi-Yuan Lin, Yu-Long Ao et al.· 0 citations
FullAttn often assigns negligible normalized mass to much of the causal score space, yet dense kernels execute the complete post-score path after forming each QK tile. We introduce MALA, a fused attention primitive that preserves score access to every legal causal interaction and uses normalized contribution to allocat...
Jing-Ze Shi, Zhang-Yang Peng, Xian-Duo Li et al.· 0 citations
CoWA, a structured attention architecture that distributes access to the causal history across KV heads, shows that full causal coverage can be a collective property of the head ensemble rather than a duplicated property of every head.
Jing-Ze Shi, Zhang-Yang Peng, Xian-Duo Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.