Skip to content

CoViST: Visual Token Compression via Composable States

Sep 2026 · 0 citations · 78 references
Computer Science

TL;DR

CoViST is proposed, a training-free framework that represents a compressed image as a composable visual state that retains its interpretation under successive reductions, enabling the same formulation to support both fixed compression before prefill and progressive compression within the decoder.

Abstract

Visual token compression lowers the inference cost of vision--language models by representing images with fewer tokens. However, most existing methods compress visual tokens to a reduced set, leaving the amount of visual evidence represented by each token and its original spatial context implicit. Therefore, the compressed representation does not explicitly encode how much visual information each representative carries or where it lies in the original image. This limitation arises even after a single reduction and becomes more pronounced when compression is repeated across decoder layers. To address this issue, we propose CoViST, a training-free framework that represents a compressed image as a composable visual state. Specifically, the state combines representative features with original positions, effective contribution weights, and reusable selection metadata. CoViST constructs this state through coverage-guided selection and conservation-based contribution composition, and explicitly incorporates its contribution and positional information into decoder attention. Each component of the state retains its interpretation under successive reductions, enabling the same formulation to support both fixed compression before prefill and progressive compression within the decoder. Experimental results on seven LLaVA-1.5-7B benchmarks show that CoViST-Fixed retains 99.9\%, 99.5\%, and 98.1\% of uncompressed performance at 192, 128, and 64 tokens, respectively, and CoViST-Pro retains 99.8\%, 99.9\%, and 99.1\% at the corresponding layer-average budgets, outperforming state-of-the-art methods under their respective budget settings. Code will be released publicly.

View source

Similar papers

Preprint Oct 2026

Foveated Compression: Selective High-Resolution Preservation for Token-Efficient VLMs

Visual tokens are a major source of inference cost in vision-language models, yet simple image downsampling remains a surprisingly strong compression baseline. This raises a complementary question: under a fixed token budget, where should visual fidelity be preserved? We introduce Foveated Compression, which encodes a...

Donghyun Han, Jangho Park, Yuseok Bae · 0 citations
Preprint Aug 2026

Aggregating Visual Information with Optimal Transport for VideoLM Token Compression

Aggregating Visual Information with Optimal Transport (AVIOT), which casts video token compression as transporting a dense empirical measure of frame observations onto a compact target measure, and adapts this construction along task and spatial axes to preserve visual information dispersed across frames under compress...

Wen-Ti Yin, Xiao-Tian Han, Jun-Yuan Shang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models

This work proposes $\delta$-Vision, which replaces repeated Transformer evolution of visual tokens with lightweight low-rank adapters that construct layer-wise visual memories while preserving all visual tokens for text retrieval, and achieves higher accuracy than visual token pruning baselines at comparable or lower c...

Jing-Di Lei, Junxian Li, Di Zhang et al. · 0 citations
#artificial intelligence Preprint Oct 2026

VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs

Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens. However, visual information is unevenly distributed: some regions require fine-grained detail, while others admit compact representations. Down...

Yuan Feng, Qi-Ze Yang, Rui-Zhe Chen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CoVisco: Codec-Native Vision Encoder with Native Token Compression for Unified Image-Video Understanding

Vision-language models face a fundamental scaling bottleneck: the number of visual tokens grows with both temporal duration and spatial resolution, making long-video understanding expensive for the vision encoder and the language model. Existing methods often compress visual tokens after dense encoding, creating a mism...

Yulong Liu, Xiao-Tian Han, Jun-Yuan Shang et al. · 0 citations
#artificial intelligence Preprint Oct 2026

VETO: Video Efficient Token Optimization for Vision Language Models

Processing long videos with Vision-Language Models (VLMs) is bottlenecked by the quadratic cost of visual tokens, making long-form inference prohibitively expensive. While single-axis compression methods mitigate this, they hit a hard efficiency floor because they treat spatial and temporal redundancy independently. We...

Gueter Josmy Faure, Hao-Ping Wang, Min-Hung Chen et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.