Preprint
Aug 2026
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models
This work proposes a two-stage adaptive token pruning strategy specifically designed for video processing that improves accuracy by +7\% on a video captioning benchmark at 10% token retention, while reducing computation TFLOPs by 95\%.
Paribesh Regmi, Qingshuang Chen, Chi Zhang et al.
· 1 citation