PCA: Persistence-Aware Compression and Aggregation for Fast Video Large Language Models
This approach substantially reduces the computation of long-context modeling, while enhancing the performance of the baseline model, achieving a speedup of 1.8$\times to 2.5$\times compared to the baseline VLLM.