WnW: Waxing-and-Waning KV Cache for Long-Form Speech LLMs
WnW (Waxing-and-Waning KV cache), which classifies KV-heads into anchor, tidal, and fixed roles via offline calibration, and preserves near-Full-Cache accuracy while keeping only 20% of audio tokens on GPU, where prefill-only baselines fail to terminate.
Yi-Ming Yao, Chenyang Lyu, Xuan-Fan Ni et al.
· 0 citations