ASPIRE: Asynchronous Batched Self-Speculative Decoding for Long-Context LLM Inference
This work proposes ASPIRE, a non-synchronized batched self-speculative decoding framework built on three components, which achieves speedup in decoding throughput over autoregressive baselines and improves average speedup by approximately $27\% over the strongest prior self-speculative baselines.