Skip to content

AdaCompVL: Adaptive Compression of Spatiotemporal and Cross-Modal Redundancy for Efficient Video-Language Learning.

Sep 2026 · IEEE Transactions on Image Processing · Vol PP · 0 citations
Medicine

Abstract

Multimodal large language models (MLLMs) have recently extended from static image understanding to video comprehension, but representing videos as frame-level token sequences incurs substantial computational overhead. Existing visual token compression methods typically rely on uniform sampling or single-dimension redundancy modeling, making them less effective across diverse scene dynamics and textual inputs. To address this issue, we propose AdaCompVL (Adaptive Compression for Video-Language), a training-free adaptive compression framework that considers both scene dynamics and semantic relevance to efficiently reduce redundant visual tokens. Specifically, we introduce a dynamic-aware switching mechanism that leverages adjacent-frame similarity to distinguish between low-and high-dynamic scenarios. For low-dynamic scenes, temporal redundancy pruning removes minimally changing tokens across frames, whereas for high-dynamic scenes, a saliency-guided spatial pruning strategy retains visually informative intra-frame tokens. To further reduce cross-modal redundancy, a mutual information-based module preserves only the visual tokens most relevant to the textual input, yielding a compact yet semantically meaningful video-language representation. Extensive experiments on multiple video-language benchmarks show that AdaCompVL removes 85% of visual tokens while retaining 99% of the original performance, achieving up to a 1.75× inference speedup. These results demonstrate that AdaCompVL produces compact and semantically relevant video representations for efficient video-language modeling. Our code is publicly available at https: //github.com/JianXin-M/AdaCompVL.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.