Skip to content

PCA: Persistence-Aware Compression and Aggregation for Fast Video Large Language Models

Jul 2026 · arXiv.org · Vol abs/2607.22726 · 0 citations · 38 references
Computer Science

TL;DR

This approach substantially reduces the computation of long-context modeling, while enhancing the performance of the baseline model, achieving a speedup of 1.8$\times to 2.5$\times compared to the baseline VLLM.

Abstract

Despite advances in Video Large Language Models (VLLMs) that have displayed promising outcomes in video understanding, the redundancy in the long-duration frames remains a hindrance to efficient reasoning. This paper introduces a training-free $\mathbf{P}$ersistence-Aware $\mathbf{C}$ompression and $\mathbf{A}$ggregation (PCA) method designed to preserve high-fidelity raw visual information before the encoding stage. PCA can be built on arbitrary VLLMs and consists of two modules: 1) A Dynamic Downsampling (DD) module that adaptively removes redundant frames by analyzing frame-wise similarity. 2) A Persistence-Aware Motion Enhancement (PAME) module that enriches each selected keyframe by aggregating the temporal context of its neighbors, ensuring that essential information is preserved even after aggressive frame reduction. Our approach substantially reduces the computation of long-context modeling, while enhancing the performance of the baseline model. Extensive experiments demonstrate that PCA consistently outperforms existing state-of-the-art approaches in both efficiency and accuracy, achieving a speedup of 1.8$\times$ to 2.5$\times$ compared to the baseline VLLM. The code is open-sourced at https://github.com/Heisenberg10110/PCA.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs

Scaling Multimodal Large Language Models (MLLMs) to long-form video understanding is bottlenecked by the explosion of visual tokens, which saturates context windows and incurs prohibitive costs. Current solutions predominantly rely on auxiliary models for token reduction but face a fundamental dilemma: lightweight enco...

Hao-Yu Guo, Yuan Feng, Junlin Lv et al. · 0 citations
Conference Aug 2026

Distilled Vision-Language Model Semantics for Edge-Deployable Content-Aware Adaptive Video Compression

With the proliferation of edge video services, content-aware adaptive video compression has become increasingly critical to balancing bandwidth efficiency and visual quality. However, conventional codec control methods mainly rely on low-level signal statistics, limiting their adaptability to high-level content semanti...

Qian Wei, En-Fang Cui, Zhi-Yuan Liang et al. · 0 citations
Preprint Aug 2026

LongVU-TTT: Causal Test-Time Training for Visual Resampling in Long Video Understanding

LongVU-TTT is introduced, which inserts a convolutional Test-Time Training (TTT) resampler with causal fast-weight updates between the vision encoder and the LLM, and is stronger than attention- and fixed-state recurrent resamplers across three benchmarks.

Mahmoud Ahmed, Sameh Abdulah, Olatunji Ruwase et al. · 0 citations
#computer vision Review Sep 2026

Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs

Video understanding has rapidly evolved toward video large language models (VideoLLMs): systems that couple video representations with pretrained large language models and condition generation on a textual prompt. Their strong performance on captioning, question answering, retrieval and temporal grounding comes at a co...

Killian Steunou, Yannis Tevissen, M. E. El Yacoubi · 0 citations
Jul 2026

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation

WaveZip is proposed, a joint signal-frequency-domain framework for efficient video inference that requires no task-specific training and can be seamlessly integrated into off-the-shelf LVLMs to boost inference efficiency.

Yuhui Zeng, Wang Chen, Jin-Fa Huang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.