Alignment Drift in Single-Model Speculative Decoding for ASR: Mechanism, Correction, and Cost
Speculative decoding speeds up generation by letting a cheap draft propose several tokens that a target model checks in one pass. In the single-model form, the draft is a lightweight module attached to the target rather than a separate model. Applying this design to Automatic Speech Recognition (ASR) introduces an extr...