DSpark-style parallel drafters have made speculative decoding highly effective, yet their draft phase remains serialized on the critical path of every round. Parallel speculative decoding (PSD) overlaps drafting with verification, yet existing methods must guess the accepted prefix and bonus token in advance: a wrong g...
Fu-Liang Liu, Xue Li, Kun Qian et al.· 0 citations
Matrix Multiplication (MatMul) faces a"generalization crisis"driven by highly dynamic tensor shapes. This crisis is particularly acute on Ascend NPUs, where explicitly controlled architectures and strict physical constraints render existing GPU-centric optimizations ineffective. To resolve this, we propose AdaptCore, a...
Yuhang Zhou, Jianglan Peng, Qian-Yu Jiang et al.· 0 citations
This paper systematically analyzes failures encountered during large-scale RL training on the Huawei Ascend platform, summarizes representative failure types, and identifies three model-side factors relevant to fault reproduction.
Yikai Wang, Chuansai Zhou, Yuhang Zhou et al.· 0 citations
Results show that template-anchored decomposition and stage-specific GPU execution can scale matrix completion beyond device-memory capacity, and show that template-anchored decomposition and stage-specific GPU execution can scale matrix completion beyond device-memory capacity.
Chengying Huan, Yu-Bo Wang, Pin-Huan Wang et al.· 0 citations
SpecLA is presented, a speculative decoding runtime for stateful linear-attention models that verifies chains and trees with topology-aware kernels, stores compact factors produced during verification to recover accepted states, and uses confidence pruning plus a target-aligned EAGLE-style drafter to feed useful candid...
Zhibin Wang, Xuying Han, Zhaohua Yang et al.· arXiv.org· 1 citation
MISA-T, a routing-layer admission policy for mixed rollout serving that combines adaptive session admission, workload-aware KV-capacity allocation, and residency-time-aware KV accounting, is presented.
Zetao Hong, Song Yuan, Yuanhao Ding et al.· 0 citations
BCE is presented, a GPU-co-designed, block-centric engine that makes range-top-k efficient by exposing a reusable intermediate representation of the data, and achieves sub-millisecond query latency and up to 308 × higher throughput than state-of-the-art GPU baselines, while performing billion-scale dynamic updates in m...
Chengying Huan, Ziheng Meng, Zhengyi Yang et al.· IEEE International Symposium...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.