Jul 2026
PCA: Persistence-Aware Compression and Aggregation for Fast Video Large Language Models
This approach substantially reduces the computation of long-context modeling, while enhancing the performance of the baseline model, achieving a speedup of 1.8$\times to 2.5$\times compared to the baseline VLLM.
Zihan Song, Shuo Ye, Bo Zhao et al.
· arXiv.org · 0 citations