H-Spec: Parallel Speculative Decoding Without a Drafter-Side KV Cache
H-Spec is a hybrid Mamba-attention parallel drafter that consumes the two target-context sources through complementary modules, and consistently achieves higher throughput while maintaining lower KV cache utilization than baselines across evaluated concurrency levels.
Wei-Fan Jiang, Krishna Teja Chitty-Venkata, Megan Flynn et al.
· 0 citations