Diffusion large language models (DLLMs) generate text through iterative block denoising, and multi-branch speculative decoding accelerates this process by verifying a main branch together with multiple draft branches in a single forward pass. While prior DLLM acceleration methods primarily exploit temporal redundancy a...
Chung-En Ho, Wei-Yu Sun, Cheng-Jhih Shih et al.· 0 citations
The authors co-design WIFiV-LPDDR, a wide-I/O LPDDR system for precision-tagged reads, BRQ-KV for a canonical low-rank-plus-INT8-residual prefix with query-dependent per-entry precision, and DAT-FFN for drift-mapped canonical replacement, adjacent-stage-corrected low-bit delta, or cached-state carry while keeping live...
Wei-Hsing Huang, Kiseok Lee, Ming-Yen Lee et al.· 0 citations
A3D-MoE addresses large language models' challenges with 3-D heterogeneous integration to improve memory bandwidth and reduce NoC overhead/energy, and a hardware resource-aware operation fusion scheduler that fuses attention/MoE operations to boost performance.
Wei-Hsing Huang, Janak Sharda, Cheng-Jhih Shih et al.· IEEE Journal on Exploratory...· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.