AdaCompVL: Adaptive Compression of Spatiotemporal and Cross-Modal Redundancy for Efficient Video-Language Learning.
Multimodal large language models (MLLMs) have recently extended from static image understanding to video comprehension, but representing videos as frame-level token sequences incurs substantial computational overhead. Existing visual token compression methods typically rely on uniform sampling or single-dimension redun...