Skip to content
Preprint

Decoupling semantics from vision: A framework for faithful visual-text compression evaluation

Aug 2026 · 1 citation · 35 references
Computer Science

TL;DR

A new evaluation framework that decouples MLLMs'capabilities to faithfully assess VTC quality is introduced, and the ZeroSense Benchmark is introduced to ensure low semantic correlation of testing samples.

Abstract

Recent visual-text compression (VTC) methods, typified by DeepSeek-OCR, report impressive high token compression ratios for long-context modeling tasks by leveraging text-to-image rendering. However, existing evaluation protocols heavily rely on downstream task performance. Such evaluation metrics fail to accurately measure text preservation due to the strong inherent linguistic priors of Multimodal Large Language Models (MLLMs). In this work, we introduce a new evaluation framework that decouples MLLMs'capabilities to faithfully assess VTC quality. Within this framework, we further introduce the ZeroSense Benchmark to ensure low semantic correlation of testing samples. By eliminating textual dependencies, our benchmark guarantees that the evaluation results are purely reflective of VTC quality, unaffected by the semantic inference capabilities of downstream models. Extensive experiments across multiple datasets demonstrate that VTC quality and downstream task accuracy diverge significantly, highlighting the necessity of our decoupled evaluation framework.

View source

Similar papers

Preprint Aug 2026

Same Semantics, Different Paths: Self-Improving Alignment for Vision-Text Compression

SPIRAL (Self-improving Path Integration and Realignment), a self-supervised alignment framework that closes this gap using only the model's own text-path behavior as supervision, requiring no external teachers or additional annotations, generalize to out-of-domain benchmarks, confirming that effective VTC hinges on ali...

Tianyu Liang, Xiangxi Zheng, Yilin Wang et al. · 2 citations
Preprint Aug 2026

VLZip: Unified Visual and Textual Compression for Interleaved Long-Context Modeling

VLZip is introduced, a framework that unifies visual and textual compression for high-fidelity reasoning within a pure Transformer, and establishes an efficient and powerful new standard for long-context multimodal AI.

Yu-Qi Zhang, Cheng Chen, Yuyu Guo et al. · 0 citations
Preprint Aug 2026

SEER: Long-Context Reasoning via Selective Visual-Text Compression

SEER is presented, a framework that learns to select query-relevant images through visual scanning and retrieve textual content only where needed, combining the efficiency of visual compression with the precision of text-based reasoning.

Jiawei Xu, Zhilin Zhai, Jinrui Fang et al. · 3 citations
Conference Aug 2026

Noun Presence Loss with Dynamic Weighting for Compositional Text-to-Image Generation

Compositional text-to-image generation requires faithful depiction of multiple objects with distinct visual attributes. While recent inference-time optimization methods have substantially improved attribute binding, entity neglect - where specified objects are absent from the generated image - remains an unresolved fai...

Trong-Tai Dam Vu, Vinh-Tiep Nguyen · 0 citations
Book Open access Aug 2026

AutoDavis: Automatic and Dynamic Evaluation Protocol of Large Vision-Language Models on Visual Question-Answering

AutoDavis is introduced, a first-of-its-kind automatic and dynamic evaluation protocol that enables on-demand benchmarking of LVLMs across specific capability dimensions and shows effectiveness and reliability, offering a new paradigm for dynamic benchmarking of multimodal intelligence.

Han Bao, Yue Huang, Yan-Bo Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs

Scaling Multimodal Large Language Models (MLLMs) to long-form video understanding is bottlenecked by the explosion of visual tokens, which saturates context windows and incurs prohibitive costs. Current solutions predominantly rely on auxiliary models for token reduction but face a fundamental dilemma: lightweight enco...

Hao-Yu Guo, Yuan Feng, Junlin Lv et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.